Data compression method and related apparatus
By performing row data permutation processing and column data division processing on table data, combined with public dictionary encoding, the problem of low compression rate of table data in the prior art is solved, and a more efficient data compression effect is achieved.
Patent Information
- Application Number
- PCT/CN2024/111695
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-12
- Filing Date
- 2024-08-13
- Publication Date
- 2025-07-17
AI Technical Summary
When the prior art compresses table data lossless data, the characteristics of table data are ignored, resulting in a low compression rate.
By performing permutation of table data based on row data, the long-tail distribution characteristics of table data are utilized, and combined with the division of column data and common dictionary encoding, the data compression rate is improved.
It effectively reduces the amount of table data moved during data compression, saves computing resources, and improves the processing efficiency and compression rate of data compression.
Smart Images

Figure CN2024111695_17072025_PF_FP_ABST
Abstract
Description
A data compression method and related device
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on January 12, 2024, with application number 202410053294.8 and application name “A data compression method and related device”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of data compression technology, and in particular to a data compression method and related devices. Background Art
[0003] With the booming development of technologies like cloud computing and big data, data is constantly being created and mined. The rapid growth in data volume and the extended storage cycles are leading to increasing data storage costs. Consequently, data compression has become a crucial research area. The emergence of various compression algorithms has not only significantly reduced storage costs across various industries, but also brought numerous additional benefits, such as improved communication efficiency across distributed systems. Data compression, also known as bitrate reduction or source coding, can help save storage space and data transmission time. Data compression can be categorized into two main types: lossless and lossy. Data compressed using lossless methods suffers no loss and can be completely restored to its pre-compression state through decompression. Lossless compression is commonly used to compress non-multimedia files, such as tables, text, source code, serialized data, and binary data, which cannot be compressed using lossy compression tools.
[0004] Lossless data compression mainly includes four types: variable-length coding, statistical compression, context coding, and dictionary coding. Dictionary coding (also known as dictionary compression) is a widely used compression method. Dictionary compression replaces recurring long strings with shorter ones.
[0005] Currently, a lossless data compression method for tabular data is as follows: the tabular data is converted into string data by column, and then encoded using a dictionary-based encoding method. However, this scheme ignores the characteristics of tabular data, resulting in a low compression rate.
[0006] Summary of the Invention
[0007] The embodiments of the present application provide a data compression method and related devices to reduce the amount of table data moved during the data compression process, save computing resources, improve the processing efficiency of data compression, and improve the data compression rate.
[0008] In a first aspect, embodiments of the present application provide a data compression method, comprising: first, a data compression device acquires first data, where the first data is tabular data consisting of multiple rows of row data and multiple columns of column data. Then, the data compression device processes the first data to generate first intermediate data, where the order of the multiple rows of data included in the first data is different from the order of the multiple rows of data included in the first data. Finally, the data compression device performs compression encoding processing on the first intermediate data to generate second data.
[0009] Exemplarily, the data compression device may be a storage device, a computing device, or a cloud server, and the embodiments of the present application are not limited thereto.
[0010] In the embodiment of the present application, based on the replacement processing of row data, the long-tail distribution characteristics of the data in the table data can be effectively utilized, and the locality of the row data in the table data can be utilized. Locality includes temporal locality and spatial locality. Temporal locality means that the data may be used multiple times in the near future, and spatial locality means that the data in the nearby storage location of the data will also be used in the near future. Therefore, the above method can reduce the amount of table data moved during the data compression process, save computing resources, save storage overhead of table data, and save computing overhead. In addition, it can effectively reduce computational complexity, improve the processing efficiency of data compression, improve the compression rate of data, and improve the compression ratio of data.
[0011] In combination with the first aspect, in a possible implementation of the first aspect, the first data is processed to generate the first intermediate data, including: mapping the row data of each row in the first data to a node to generate a node set, the node set including multiple nodes; processing the first data according to the connection order of the multiple nodes included in the node set to generate the first intermediate data, and the order of the multiple row data of the first intermediate data is the same as the connection order of the multiple nodes included in the node set.
[0012] Optionally, the node set sequentially connects multiple nodes to form a path based on the locality of the rows of data corresponding to the multiple nodes. The path passes through each node in the node set only once. The connection order of the multiple nodes included in the node set refers to the connection order of the nodes in the path.
[0013] It should be noted that a node in the node set can also be obtained by mapping multiple rows of data in the first intermediate data, and this embodiment of the present application does not limit this.
[0014] In the embodiment of the present application, performing substitution processing on the first data according to the above substitution rule to generate the first intermediate data can effectively utilize the locality of the table data and improve the data compression rate.
[0015] In combination with the first aspect, in a possible implementation manner of the first aspect, a connection order of the multiple nodes included in the node set is the same as an order indicated by a shortest Hamiltonian path connecting the nodes in the node set.
[0016] A Hamiltonian path is a path that starts at a point in a graph and follows the edges, passing through every point in the graph exactly once. The shortest Hamiltonian path is the shortest path that passes through every point in the graph exactly once.
[0017] In the embodiments of the present application, the shortest Hamiltonian path connecting each node in the node set refers to the optimal solution for sequentially connecting multiple nodes in the node set. The first data is permuted according to the order indicated by the shortest Hamiltonian path connecting each node in the node set, so that the data compression ratio of the permuted first intermediate data is optimized, effectively utilizing the locality of the table data and effectively improving the data compression rate.
[0018] In conjunction with the first aspect, in one possible implementation of the first aspect, the first intermediate data further includes first indication information, the first indication information being used to indicate the order of each row of data in the first intermediate data within the first data. Exemplarily, the first indication information is the sequence number of the row of data in the first data. The first indication information guides the data decompression process to restore the first data based on the first intermediate data, thereby improving the data decompression processing rate.
[0019] In combination with the first aspect, in a possible implementation of the first aspect, the first intermediate data is compressed and encoded to generate the second data, including: processing the first intermediate data to obtain second intermediate data, the second intermediate data at least including first column data and second column data, the first column data including one or more columns of column data of the first intermediate data, and the second column data including one or more columns of column data of the first intermediate data; compressing and encoding the second intermediate data to generate the second data.
[0020] In the embodiments of the present application, partitioning and joining processing refers to partitioning multiple columns of original table data into one or more column data sets, each of which includes one or more columns of data, and then joining the one or more columns of data included in the column data sets into one column of data, so that the number of columns of data in the new table data obtained after the partitioning and joining processing is less than the number of columns of data in the original table data. Partitioning and joining processing is performed on the table data, utilizing the correlation between the columns of data in the table data to remove redundancy between the columns of the table data, simplify the dimensionality of the table data, reduce computational complexity, and effectively improve the processing efficiency of data compression.
[0021] In conjunction with the first aspect, in a possible implementation of the first aspect, the second intermediate data further includes second indication information, where the second indication information is used to indicate a correspondence between column data of the second intermediate data and column data of the first intermediate data. The first indication information guides the data decompression process to restore the first intermediate data based on the second intermediate data, thereby improving the data decompression processing rate.
[0022] In conjunction with the first aspect, in a possible implementation of the first aspect, the first data is processed to generate fifth intermediate data, the fifth intermediate data including at least third column data and fourth column data, the third column data including one or more columns of column data of the first data, and the fourth column data including one or more columns of column data of the first data; and the fifth intermediate data is processed to generate sixth intermediate data, the sixth intermediate data including multiple rows of data in an order different from the order of the multiple rows of data included in the first intermediate data. This improves the implementation flexibility of the solution.
[0023] In combination with the first aspect, in a possible implementation of the first aspect, the second intermediate data is compressed and encoded to generate the second data, including: using a first common dictionary to encode the second intermediate data to obtain third intermediate data, the first common dictionary is constructed by multiple second intermediate data, and the multiple second intermediate data correspond to multiple first data; and compressing and encoding the third intermediate data to generate the second data.
[0024] It should be noted that, in the embodiments of the present application, a variety of different public dictionaries can be constructed for different data files to be encoded (or referred to as encoding objects). In another example, the first data is encoded using a second public dictionary, which is constructed from multiple first data obtained offline. In another example, the first intermediate data is encoded using a third public dictionary, which is constructed from multiple first intermediate data obtained offline. Using a public dictionary for encoding improves the data compression rate while saving dictionary storage overhead.
[0025] In combination with the first aspect, in a possible implementation of the first aspect, the first data is replaced according to the granularity of row data to generate the first intermediate data, including: using a second common dictionary to encode the first data to obtain fourth intermediate data, the second common dictionary is constructed by multiple first data; processing the fourth intermediate data to generate the first intermediate data, the order of the multiple row data included in the first intermediate data is different from the order of the multiple row data included in the fourth intermediate data.
[0026] The data compression method of the embodiment of the present application does not restrict the order in which the permutation processing based on the granularity of row data, the partitioning and joint processing based on the granularity of column data, and the encoding processing using a common dictionary are executed. The data compression method of the embodiment of the present application can also adopt any one or more of the above processing methods according to actual needs.
[0027] In combination with the first aspect, in a possible implementation of the first aspect, the first data is permuted at the granularity of row data to generate the first intermediate data, including: mapping the row data of each row in the first data to a node to generate a node set, the node set includes multiple nodes corresponding to the row data of the multiple rows included in the first data; sorting the node set to generate an index sequence, the index sequence includes a series of ordered index values, and each index value in the index sequence corresponds to a node in the node set; permuting the first data at the granularity of row data according to the index sequence to generate the first intermediate data composed of sorted row data of the multiple rows, and the sorting of the multiple row data included in the first intermediate data is the same as the sorting of the index sequence.
[0028] In the embodiment of the present application, the permutation optimization problem of row data is converted into an asymmetric TSP problem, and an index sequence guiding the permutation processing of row data is obtained through heuristic solution, thereby effectively improving the processing efficiency of data compression.
[0029] In combination with the first aspect, in a possible implementation of the first aspect, the node set is sorted to generate the index sequence, including: forming a directed complete graph based on the node set, wherein any two nodes in the node set in the directed complete graph are connected by an edge; solving the shortest Hamiltonian path of the directed complete graph; and determining the index sequence based on the shortest Hamiltonian path of the directed complete graph, wherein the order of the index sequence is consistent with the order of the index values of the nodes in the shortest Hamiltonian path of the directed complete graph.
[0030] In the embodiments of the present application, the shortest Hamiltonian path of a directed complete graph indicates the correlation between nodes in a node set, that is, the correlation between rows of data in a table. Therefore, by permuting rows of table data based on the shortest Hamiltonian path of the directed complete graph, similar rows of data can be shifted to adjacent positions in the table, thereby emphasizing the locality of the table data and improving the data compression rate.
[0031] In combination with the first aspect, in a possible implementation of the first aspect, solving the shortest Hamiltonian path of the directed complete graph includes: determining the weight of the edge connecting any two nodes in the directed complete graph, wherein the weight of the edge is obtained by subtracting the second function value from the first function value, the second function value is the number of consecutively repeated elements in the first data, the first function value is the number of consecutively repeated elements in the first permuted data, and the first permuted data is the tabular data obtained after permuting the row data corresponding to the two nodes constituting the edge in the first data; step (1), initializing the feasible edge set and the target edge set, the initialized target edge set is empty, when any node in the directed complete graph When a subgraph composed of one or more of the edges does not include a Hamiltonian cycle, the edges constituting the subgraph are used as the feasible edges, and the initialized feasible edge set includes all the feasible edges in the directed complete graph; step (2): randomly selecting one of the two edges with the smallest weight in the feasible edge set to generate the target edge set; step (3): deleting the edges included in the target edge set from the feasible edge set; repeating steps (2) and (3) until the subgraph composed of the target edge set includes the shortest Hamiltonian path of the directed complete graph, and the shortest Hamiltonian path of the directed complete graph is a path connecting all nodes in the node set.
[0032] In the embodiment of the present application, a greedy algorithm is used to solve the shortest Hamiltonian path of a node set, which reduces the computational complexity, improves the encoding speed, and improves the compression speed of data compression.
[0033] In conjunction with the first aspect, in a possible implementation of the first aspect, processing the first intermediate data to obtain the second intermediate data includes: determining, based on a dynamic programming algorithm, an optimal solution for partitioning and jointly processing the column data combination of the first intermediate data according to the granularity of the column data, the optimal solution including at least a first column joint and a second column joint, the first column joint including one or more schemes for partitioning and jointly processing the column data of the first intermediate data, the second column joint including one or more schemes for partitioning and jointly processing the column data of the first intermediate data, and the intersection between the column data corresponding to the first column joint and the column data corresponding to the second column joint is empty; determining, based on the optimal solution, the first column data and the second column data of the second intermediate data, wherein the first column data corresponds to the first column joint in the optimal solution, the second column data corresponds to the second column joint in the optimal solution, the first column data includes one or more columns of column data, and the second column data includes one or more columns of column data. The first column data corresponds to the partitioning and joint scheme included in the first column joint, the second column data corresponds to the partitioning and joint scheme included in the second column joint, the first column data includes one or more columns of column data, and the second column data includes one or more columns of column data.
[0034] It is understandable that the optimal solution may include more column join schemes, such as a third column join, a fourth column join, and a fifth column join, etc. Each column join included in the optimal solution includes one or more schemes for partitioning and joining column data. Based on the column join, it can be determined that one or more columns of column data of the first intermediate data are partitioned and joined to obtain one or more columns of new column data. The new column data obtained through the partitioning and joining includes the column data of one or more columns in the first intermediate data.
[0035] In the embodiment of the present application, a dynamic programming algorithm is used to solve the partitioning and combining scheme of column data, which reduces the computational complexity, can improve the encoding speed, and improve the compression speed of data compression.
[0036] In combination with the first aspect, in a possible implementation of the first aspect, the optimal solution satisfies: the sum of the size of the first compressed data and the size of the second compressed data is the minimum value of the data after the first intermediate data is arbitrarily divided and jointly processed and subjected to data compression processing, the first compressed data is the data obtained by data compression processing of the column data included in the first column union, and the second compressed data is the data obtained by data compression processing of the column data included in the second column union.
[0037] In conjunction with the first aspect, in one possible implementation of the first aspect, the optimal solution for performing partitioning and joint processing on the first intermediate data may be an optimal solution generated by a data compression device executing the data compression method based on one or more offline first intermediate data. When the first data is compressed and encoded online, the first intermediate data (or first data) generated by the first data is directly partitioned and joint processed using the offline-generated optimal solution to generate second intermediate data.
[0038] In combination with the first aspect, in a possible implementation of the first aspect, the method also includes: obtaining multiple second intermediate data; constructing the first common dictionary based on the multiple second intermediate data, the first common dictionary including the data to be matched and the coding values corresponding to the data to be matched.
[0039] In the embodiment of the present application, data encoding based on a common dictionary can effectively utilize the correlation between tables, solve the problem that a single data in some columns of a table occupies too much memory, and achieve compact representation of the data.
[0040] In combination with the first aspect, in a possible implementation manner of the first aspect, the to-be-matched data corresponds to column data of any one or more columns in the second intermediate data.
[0041] It can be understood that the data to be matched may also correspond to one or more rows of row data in the second intermediate data, or the data to be matched may also correspond to any portion of data in the second intermediate data.
[0042] In combination with the first aspect, in a possible implementation of the first aspect, constructing the first common dictionary based on the multiple second intermediate data includes: generating a dictionary table based on the multiple second intermediate data, the dictionary table including data of the multiple second intermediate data; dividing the multiple data to be matched in the dictionary table into one or more data sets to be matched based on the statistical data of the data to be matched in the dictionary table, each data set to be matched includes at least one data to be matched, and the statistical data indicates the total amount of the data to be matched in the dictionary table; encoding the data to be matched to generate a coding value corresponding to the data to be matched; constructing one or more first common dictionaries based on the data to be matched and the coding value corresponding to the data to be matched, wherein each first common dictionary includes one data set to be matched.
[0043] In the embodiment of the present application, the above method is used to reduce the overhead of querying the dictionary, save encoding time, and improve the efficiency of compressing data.
[0044] In combination with the first aspect, in a possible implementation manner of the first aspect, the encoding process includes any one of the following: Huffman coding, entropy coding, or arithmetic coding.
[0045] In combination with the first aspect, in a possible implementation manner of the first aspect, the method further includes: acquiring multiple second intermediate data; and updating the first public dictionary according to distribution changes of data in the multiple second intermediate data.
[0046] In this embodiment of the present application, the applicability period of the first common dictionary is determined by detecting changes in data distribution. If the first common dictionary is no longer applicable to the current data distribution, a new first common dictionary is established to adapt to data changes and improve data compression.
[0047] In combination with the first aspect, in a possible implementation of the first aspect, performing compression encoding processing on the first intermediate data to generate the second data includes: performing run-length encoding processing on the first intermediate data to generate the second data.
[0048] It is understood that compression coding methods include but are not limited to: run-length coding, Huffman coding, differential coding, LZ77 coding, or Shannon-Fano coding, etc., to further improve the data compression rate.
[0049] In a second aspect, embodiments of the present application provide a data decompression method. The data decompression method provided in the second aspect is obtained by parsing the data compression method in the first aspect and any one of the data compression methods in the first aspect. For example, the decompression device reversely executes the data compression method in the first aspect and any one of the data compression methods in the first aspect.
[0050] In one possible implementation, the method includes: a decompression device obtains second data, the second data includes tabular data consisting of multiple rows of row data and multiple columns of column data; the decompression device decompresses the second data to generate first intermediate data; the decompression device processes the first intermediate data to generate first data, and the order of the multiple row data included in the first intermediate data is different from the order of the multiple row data included in the first data.
[0051] Exemplarily, the decompression device may be a storage device, a computing device, or a cloud server, and the embodiments of the present application are not limited thereto.
[0052] In the embodiment of the present application, the above-mentioned data decompression method can improve the decoding speed and the data decompression speed.
[0053] In one possible implementation, processing the first intermediate data to generate the first data includes: obtaining first indication information from the first intermediate data, the first indication information being used to indicate the order of each row of data in the first intermediate data within the first data; and processing the first intermediate data based on the first indication information to generate the first data. In conjunction with the first indication information, the order of the rows of data in the first intermediate data is restored, thereby ensuring the accuracy of data decompression, improving processing speed, and further improving the speed of data decompression.
[0054] In one possible implementation, decompressing the second data to generate the first intermediate data includes: decompressing the second data to generate second intermediate data, the second intermediate data including at least first column data and second column data; and processing the second intermediate data to generate the first intermediate data, wherein one or more columns of data in the first intermediate data are derived from the first column data, and one or more columns of data in the first intermediate data are derived from the second column data. Using this data decompression method can improve decoding speed and data decompression speed.
[0055] In one possible implementation, processing the second intermediate data to generate the first intermediate data includes: obtaining second indication information from the second intermediate data, the second indication information being used to indicate a correspondence between column data of the second intermediate data and column data of the first intermediate data; and processing the second intermediate data based on the second indication information to generate the first intermediate data. In combination with the second indication information, the column data of the second intermediate data is split and restored to the first intermediate data, thereby improving processing speed, ensuring data decompression accuracy, and further improving data decompression speed.
[0056] In one possible implementation, decompressing the second data to generate the second intermediate data includes: decompressing the second data to generate third intermediate data; and decoding the third intermediate data using a first common dictionary to obtain the second intermediate data, wherein the first common dictionary is constructed from a plurality of the second intermediate data, and the plurality of the second intermediate data corresponds to a plurality of the first data. Using this data decompression method can improve decoding speed and data decompression speed.
[0057] In one possible implementation, processing the first intermediate data to generate the first data includes: processing the first intermediate data to obtain fourth intermediate data, wherein the order of the plurality of rows of data included in the first intermediate data is different from the order of the plurality of rows of data included in the fourth intermediate data; and decoding the fourth intermediate data using a second common dictionary to obtain the first data, wherein the second common dictionary is constructed from the plurality of first data. Using the above data decompression method can improve decoding speed and data decompression speed.
[0058] In one possible implementation, decompressing the second data to generate the first intermediate data includes performing run-length encoding decoding on the second data to generate the first intermediate data. Using the above data decompression method can increase decoding speed and data decompression speed.
[0059] In a third aspect, an embodiment of the present application proposes a data compression device, which includes a processing unit and a transceiver unit. The data compression device is used to execute the first aspect and any one of the methods described in the first aspect.
[0060] In a fourth aspect, an embodiment of the present application proposes a decompression device, which includes a processing unit and a transceiver unit, and is used to execute the method described in the aforementioned second aspect and any one of the second aspects.
[0061] In a fifth aspect, an embodiment of the present application provides a chip, which includes an interface circuit and a processing circuit. The interface circuit and the processing circuit are interconnected through lines, and the processing circuit is used to run computer programs or instructions to perform the method of the first aspect or the second aspect.
[0062] Optionally, the chip includes at least one processor and a communication interface, the communication interface and the at least one processor are interconnected via a line, and the at least one processor is used to run a computer program or instruction to perform the method of the first aspect or the second aspect.
[0063] Optionally, the communication interface of the core chip can be an input / output interface, a pin or a circuit, etc.
[0064] In conjunction with the fifth aspect, in a first implementation of the fifth aspect of the embodiments of the present application, the chip described above in the present application further includes at least one memory, wherein the at least one memory stores instructions. The memory can be a storage unit within the chip, such as a register, a cache, etc., or can be a storage unit of the chip (such as a read-only memory, a random access memory, etc.).
[0065] In a sixth aspect of an embodiment of the present application, a storage device is provided, comprising at least one processor coupled to a memory; the memory is used to store programs or instructions; and the at least one processor is used to execute the programs or instructions so that the device can implement any possible implementation method of the aforementioned first aspect or second aspect.
[0066] The seventh aspect of an embodiment of the present application provides a storage device, including a communication interface for inputting and / or outputting signaling or data; and a processor for executing a computer-executable program so that the device can implement any possible implementation method of the aforementioned first aspect or second aspect.
[0067] In an eighth aspect of an embodiment of the present application, a storage device is provided, comprising at least one logic circuit and an input / output interface; the input / output interface is used to input or output information; and the logic circuit is used to execute any possible implementation method as described in the first or second aspect above.
[0068] A ninth aspect of the present application provides a storage system, comprising a storage device as described in any implementation of the sixth aspect above.
[0069] In a tenth aspect, the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer-readable storage medium is run on a computer, the computer executes the method described in the first or second aspect above.
[0070] In an eleventh aspect, the present application provides a computer program product, which, when executed on a computer, enables the computer to execute the method described in the first or second aspect above.
[0071] The twelfth aspect of the present application provides a computing system, which includes an encoding device and / or a decompression device, wherein the encoding device is used to execute the method as described in any one of the first aspects, and the decompression device is used to execute the method as described in any one of the second aspects.
[0072] The thirteenth aspect of the present application provides a storage system, which includes an encoding device and / or a decompression device, wherein the encoding device is used to execute the method as described in any one of the first aspects, and the decompression device is used to execute the method as described in any one of the second aspects. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] FIG1 is a flow chart of a dictionary-based lossless compression encoding method;
[0074] FIG2 is a schematic diagram of the structure of a storage device provided in an embodiment of the present application;
[0075] FIG3 is a flow chart of an embodiment of a data compression method according to an embodiment of the present application;
[0076] Figure 4 is a schematic diagram of the ATSP problem with 4 nodes;
[0077] FIG5 is a schematic diagram of a replacement process in an embodiment of the present application;
[0078] FIG6 is a schematic flow chart of a replacement processing method in an embodiment of the present application;
[0079] FIG7 is a schematic diagram of mapping inline data to nodes according to an embodiment of the present application;
[0080] FIG8 is a schematic diagram of a directed complete graph in an embodiment of the present application;
[0081] FIG9 is a schematic diagram of solving the shortest Hamiltonian path in an embodiment of the present application;
[0082] FIG10 is a flow chart of a partitioning and joint processing method according to an embodiment of the present application;
[0083] FIG11 is a schematic diagram of division and joint processing in an embodiment of the present application;
[0084] FIG12 is a flow chart of an encoding method based on a public dictionary in an embodiment of the present application;
[0085] FIG13 is a schematic diagram of an encoding method based on a public dictionary in an embodiment of the present application;
[0086] FIG14 is a schematic diagram of compression encoding processing of third intermediate data in an embodiment of the present application;
[0087] FIG15 is a flow chart of a method for constructing a first public dictionary according to an embodiment of the present application;
[0088] FIG16a is a schematic diagram of constructing a first public dictionary in an embodiment of the present application;
[0089] FIG16b is a schematic diagram of an application scenario in an embodiment of the present application;
[0090] FIG17 is a diagram showing data frequency statistics of a dictionary table in an embodiment of the present application;
[0091] FIG18 is a schematic diagram of a flow chart of a data decompression method according to an embodiment of the present application;
[0092] FIG19a is a schematic diagram of an application scenario in an embodiment of the present application;
[0093] FIG19b is a schematic diagram of an application scenario in an embodiment of the present application;
[0094] FIG20 is a schematic structural diagram of a data compression device provided in an embodiment of the present application;
[0095] FIG21 is a schematic structural diagram of a decompression device provided in an embodiment of the present application;
[0096] FIG22 is a schematic structural diagram of a storage device 2201 provided in an embodiment of the present application. DETAILED DESCRIPTION
[0097] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all of the embodiments. The terms "first", "second" and corresponding terminology labels in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way are interchangeable where appropriate, and this is merely a way of distinguishing objects of the same properties when describing the embodiments of the present application. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, so that a process, method, system, product or device that includes a series of units is not necessarily limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or devices.
[0098] In the description of this application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in this application is merely a description of the association relationship of associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of this application, "at least one" refers to one or more items, and "multiple items" refers to two or more items. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.
[0099] The following are some technical concepts involved in this application:
[0100] Data compression: also known as bit rate reduction technology or source coding, specifically the process of representing information with fewer data bits (or other information-related units) than in the unencoded state according to a specific coding mechanism. Data compression can be achieved through a data compression algorithm. Depending on the coding effect, data compression algorithms can be divided into lossless data compression and lossy data compression. Data compression (or compression) in the embodiments of the present application can also be referred to as data encoding (or encoding), and correspondingly, data decompression (or decompression) can also be referred to as data decoding (or decoding or decoding), which is not limited in the embodiments of the present application.
[0101] Lossless data compression: A compression method in which compressed data can be completely restored to its pre-compression state through decompression. Lossless data compression is typically used in business scenarios where reliability is paramount. For example, it can be used to compress non-multimedia files such as text, source code, serialized data, and binary files that cannot be compressed using lossy compression tools. Lossless data compression can be achieved through dictionary encoding. Dictionary encoding is essentially a preprocessing process on the data stream. Essentially, it replaces recurring long strings with shorter strings to achieve compression.
[0102] Tabular data: Data presented in tabular form can be presented in either wide or narrow table form. A wide table, in the literal sense, is a database table with many fields. A wide table usually refers to a database table that associates indicators, dimensions, and attributes related to a business topic. For example, Table 1 below shows tabular data in the telecommunications domain.
[0103] Table 1
[0104] Tabular data represents all sample data presented in the table. A row is a sample, and a column is a feature. For example, in Table 1, the row containing the user named Zhang San is a sample, which includes Zhang San's mobile phone number, the location of the mobile phone number, and the package type of the mobile phone number. Column elements such as location and package type in Table 1 are features, while location A, location B, package 1, and package 2 can be called the corresponding feature values.
[0105] As can be seen from Table 1, tabular data can have one or more object description features. Each object description feature has semantics, which means that the semantics are given meaning. In other words, each object description feature has a specific meaning. For example, the object description feature of location is used to indicate the country or city to which a mobile phone number belongs.
[0106] Spatial Locality: Information that will be used in the future is likely to be close to the information currently being used in terms of spatial address. The above information has spatial locality.
[0107] A common dictionary-based lossless data compression coding method is shown in Figure 1, which is a flow chart of a dictionary-based lossless compression coding method. The method converts table data into string data in the form of column data concatenation (referred to as column concatenation), and then uses dictionary-based coding. Specifically, the encoder that runs the lossless compression coding method includes a sliding window (or a dynamic window) and a pre-read buffer. The pre-read buffer corresponds to the sliding window, and the pre-read buffer is used to cache the encoded data. The longest matching data with the encoded data is searched in the sliding window (the sliding window includes the data to be encoded). Then, the longest matching data is represented as a triple: matching distance, matching length (or the length of the longest match) and the next character to achieve data compression coding.
[0108] However, the above data compression method converts table data into string data in a column-concatenated manner, ignoring the spatial locality of table data, resulting in a low compression rate.
[0109] Based on the above problems, an embodiment of the present application proposes a data compression method. After obtaining first data, the first data is processed to generate first intermediate data. The order of the multiple row data included in the first intermediate data is different from the order of the multiple row data included in the first data. The first data is tabular data consisting of multiple rows of row data and multiple columns of column data. The first intermediate data is then compressed and encoded to generate second data. After obtaining the first data, the first data is subjected to row data-based permutation processing considering the locality of the tabular data. This can effectively utilize the long-tail distribution characteristics of the column data in the tabular data and improve the data compression rate.
[0110] The following describes an embodiment of the present application with reference to the accompanying drawings. First, the hardware structure of the storage device in the embodiment of the present application will be illustrated. It will be appreciated that the storage device may also be a computing device with storage functionality, and this is not a limitation of the present embodiment. The storage device may be used to execute the data compression method in the embodiment of the present application, for example, the storage device may execute the functions associated with the data compression device in the embodiment of the present application. The storage device may also be used to execute the data decompression method corresponding to the data compression method in the embodiment of the present application, for example, the storage device may execute the functions associated with the decompression device in the embodiment of the present application. Figure 2 is a schematic diagram of the structure of a storage device provided in the embodiment of the present application. As shown in Figure 2, the storage device 100 is connected to the application server 200 via a switch 300. The application server 200 is a computer that runs application programs. The application server 200 may be a physical machine or a virtual machine. Physical machines include, but are not limited to, desktop computers, servers, laptops, and mobile devices. The application server 200 accesses the storage device 100 via the switch 300 to access data. However, the switch 300 is an optional device; the application server 200 may also communicate directly with the storage device 100 over a network.
[0111] The storage device 100 shown in Figure 2 is a centralized storage device. A centralized storage device is characterized by a unified entry point through which all data from external devices must pass. This entry point is the centralized storage device's engine 120. Engine 120 is the core component of the centralized storage device, implementing many of the device's advanced functions.
[0112] As shown in Figure 2, there are one or more controllers in the engine 120. Figure 2 uses the example of an engine 120 containing two controllers for explanation. There is a mirror channel between controller 0 and controller 1. Then, after controller 0 writes a copy of data into its memory 124, it can send a copy of the data to controller 1 through the mirror channel, and controller 1 stores the copy in its own local memory 124. In this way, controller 0 and controller 1 back up each other. When controller 0 fails, controller 1 can take over the business of controller 0. When controller 1 fails, controller 0 can take over the business of controller 1, thereby avoiding the unavailability of the entire storage device due to hardware failure. When there are 4 controllers deployed in the engine 120, there is a mirror channel between any two controllers, so any two controllers back up each other.
[0113] The engine 120 also includes a front-end interface 125 and a back-end interface 126. The front-end interface 125 is used to communicate with the application server 200, thereby providing storage services for the application server 200. The back-end interface 126 is used to communicate with the hard disk 140 to expand the capacity of the storage device 100. Through the back-end interface 126, the engine 120 can connect to more hard disks 140, thereby forming a very large storage resource pool.
[0114] In terms of hardware, the controller 0 includes at least a processor 123 and a memory 124. The processor 123 can be a central processing unit (CPU), which is used to process data access requests from outside the storage device 100 (such as the application server 200 or other storage devices), and also to process requests generated within the storage device 100. For example, when the processor 123 receives write data requests sent by the application server 200 through the front-end port 125, it temporarily stores the data streams in these write data requests in the memory 124. When the total amount of data in the memory 124 reaches a certain threshold, the processor 123 compresses the data streams stored in the memory 124 and sends them to the hard disk 140 for persistent storage through the back-end port.
[0115] Memory 124 refers to internal storage that directly exchanges data with processor 123. It can read and write data at any time and at high speed, serving as temporary data storage for the operating system or other running programs. Memory includes at least two types of memory. For example, memory can be either random access memory (RAM) or read-only memory (ROM). For example, RAM is dynamic random access memory (DRAM) or storage class memory (SCM); ROM can be programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), etc. In actual applications, controller 0 can be configured with multiple memories 124, as well as memories of different types. This embodiment does not limit the number and type of memories 124. In addition, memory 124 can be configured to have a power-saving function. This power-saving function ensures that data stored in memory 124 will not be lost even if the system loses power and then powers on again. Memory with a power-saving function is called non-volatile memory.
[0116] The hardware components and software structure of controller 1, as well as other controllers not shown in FIG2 , are similar to those of controller 0 and are not further described here. It should be noted that FIG2 shows only one engine 120 . However, in actual applications, the storage system may include two or more engines 120 , and multiple engines 120 may be used for redundancy or load balancing.
[0117] The storage device 100 shown in FIG2 is a centralized storage device with integrated disk and controller. In this device, the engine 120 has a hard disk slot, and the hard disk 140 can be directly deployed in the engine 120. The back-end interface 126 is an optional configuration. When the storage space of the storage device 100 is insufficient, more hard disks or hard disk enclosures can be connected through the back-end interface 126. In other possible implementations of the embodiments of the present application, the storage device 100 can also be a centralized storage device with separate disk and controller, or a storage device in a distributed storage system.
[0118] FIG2 provides a detailed description of the hardware structure of a storage device 100 for executing the data compression method. Next, the data compression method provided by the embodiment of the present application will be described in detail from the perspective of the storage device 100 in conjunction with the accompanying drawings. Please refer to FIG3 , which is a schematic flow chart of an embodiment of the data compression method of the present application. The data compression method provided by the embodiment of the present application includes:
[0119] S1. Acquire first data, where the first data is table data consisting of a plurality of rows of row data and a plurality of columns of column data.
[0120] The first data acquired in step S1 is a tabular data consisting of a plurality of rows of row data and a plurality of columns of column data. The embodiment of the present application does not limit the type of the first data, and the first data can be any tabular data.
[0121] In one example, the original data is shown in Table 1.
[0122] In another example, the original first data is shown in Table 2.
[0123] Table 2
[0124] In another example, the first data is shown in Table 3. In Table 3, the second column is the Internet Protocol (IP) address of the device, the third column is the name of the device (deviceName), the fourth column is the source label switching path (srclsp), the fifth column is the source autonomous domain system (srcAS), and the sixth column is bits per second (bps).
[0125] Table 3
[0126] S2. Perform replacement processing on the first data according to the granularity of row data to generate first intermediate data.
[0127] In step S2, considering the long-tail distribution characteristics of the table data, the entire row data of the table data is replaced to highlight the locality of the data. The long-tail distribution is also called the power-law distribution. In the table data, the long-tail distribution characteristics are reflected in the fact that the elements (i.e., data) with extremely high frequency of occurrence account for an extremely low proportion in the total data, while the elements with extremely low frequency of occurrence account for an extremely high proportion in the total data. By utilizing the long-tail distribution characteristics of the table data, the table data is replaced according to the granularity of the row data, with the aim of placing the same row data in the same area of the table. In the subsequent compression coding process, such as the run-length coding process, multiple identical row data can be encoded into coding values with smaller storage space, thereby effectively improving the data compression rate.
[0128] The first data is processed to generate first intermediate data, and the order of the multiple row data included in the first intermediate data is different from the order of the multiple row data included in the first data. For the sake of convenience, the above-mentioned processing of the first data can be called replacement processing (or row replacement processing), and replacement processing refers to replacing the row data in the tabular data according to the granularity of the row data. For example, as shown in Table 4 and Table 5, Table 4 is the data before the replacement processing, that is, the first data. The first data shown in Table 4 is replaced, and the generated first intermediate data is shown in Table 5. Specifically, the row data of sequence number 1 in Table 4 is replaced with the row data of sequence number 5 to obtain the first intermediate data shown in Table 5.
[0129] Table 4
[0130] Table 5
[0131] It should be noted that other processing can be performed before the first data is replaced in step S2. In other words, in step S2, the intermediate data obtained through other processing can be replaced according to the granularity of row data to generate first intermediate data. This embodiment of the present application does not limit this.
[0132] Optionally, the first intermediate data may further include first indication information, which is used to indicate the order of each row of data in the first intermediate data within the first data. For example, the first indication information is the "sequence number" in the first column of the data in Table 5. It is understood that the first indication information may also be independent of the first intermediate data, for example, by performing permutation processing on the first data and outputting the first intermediate data and the first indication information.
[0133] The following describes the replacement processing algorithm:
[0134] Taking the table data before replacement as the first data as an example, the goal of performing replacement processing on the first data at the granularity of row data is: in a certain replacement scheme, the file size of the table data after replacement processing and then after data compression processing has the best compression ratio compared to any other replacement scheme (or the smallest file size after data compression, or the smallest data volume after data compression). The optimized replacement scheme (or the mathematical optimization model of the replacement scheme, or the optimal solution of the replacement scheme) can be expressed as the following formula:
[0135] Wherein, π* is the file size of the compressed output of the table data in the optimal replacement scheme. The table data before replacement (for example, the first data) is a table data including n rows of row data and m columns of column data. The table data before replacement can be expressed as T n,m , n is a positive integer greater than 1, m is a positive integer greater than 1. T [i,:] represents the i-th row of table T, T [:,j] represents the j-th column of table T, i is an integer greater than or equal to 0 and less than or equal to n, and j is an integer greater than or equal to 0 and less than or equal to m. π represents the permutation process for the rows in table T. It represents the πth in the table data [1] The data of row j column, πth [1] Row refers to the first row of table T permuted to π [1] C(T [:,j] ) represents T [:,j] The size of the output file after compression encoding.
[0136] In a possible implementation, the mathematical optimization model of the above replacement scheme can be mapped to an asymmetric traveling salesman problem (ATSP) for solution. Specifically, each row of the table data before replacement is mapped to a node of the graph, such as the T [i,:] The rows are mapped to nodes i, and the multiple nodes mapped to the table data before the permutation are called node sets. Then, each node in the node set is connected to other nodes through edges to form a directed complete graph. In this directed complete graph, any two nodes are connected by edges, and each edge corresponds to a weight (or cost). For example, the weights corresponding to nodes i1 and i2 are called: Node i1 is mapped from row i1 of table T, and node i2 is mapped from row i2 of table T. It can be expressed as the following formula:
[0137] Among them, CntDup is a function that counts the number of consecutively repeated elements (or characters or data). [:,j] ) indicates T [:,j] The number of consecutively repeated elements in , It means that row i2 is moved between row i1 and row i1+1 after replacement. ij ≠c ji For ease of understanding, please refer to Figure 4 for an example, which is a schematic diagram of the ATSP problem with 4 nodes. [1,:] Node T [2,:] Node T [3,:] and node T [4,:] For example, the above four nodes form a directed complete graph, and the cost between any two nodes is shown in Figure 4. The optimal solution to the ATSP problem is the shortest Hamiltonian path connecting each node in the graph.
[0138] In summary, the optimization objective of the mathematical optimization model of the replacement scheme is as follows:
[0139] The above ATSP problem can be solved by a greedy algorithm (a greedy algorithm can also be called a greedy solution algorithm). That is, the directed complete graph including multiple nodes corresponding to the above ATSP problem can be solved by a greedy algorithm to obtain the optimal solution (or approximate solution). (1) Determine the feasible edges in the directed complete graph. A feasible edge means that after the edge is added to the target edge set, the subgraph composed of one or more edges included in the target edge set does not include a Hamiltonian cycle. In this case, the edge is called a feasible edge. (2) Sort the feasible edges in the feasible edge set according to their weights. The initialized feasible edge set includes all feasible edges in the directed complete graph. (3) Randomly select one or more feasible edges with the smallest weight in the feasible edge set and add it to the target edge set. Then delete the selected feasible edge from the feasible edge set. (4) Repeat the above step (3) until the target edge set forms a shortest Hamiltonian path including {1,…,n}. The shortest Hamiltonian path including {1,…,n} refers to the shortest Hamiltonian path connecting all nodes in the directed complete graph. The connection order of the nodes in the shortest Hamiltonian path is the optimal solution of the permutation scheme of the table data T.
[0140] For ease of understanding, the above-mentioned replacement processing method is described below in conjunction with the accompanying drawings. Please refer to Figure 5, which is a schematic diagram of a replacement processing in an embodiment of the present application. Take the first data before replacement processing as a tabular data including row data of the zeroth row, the first row, the second row and the third row and column data of the Ath column, the Bth column, the Cth column, the Dth column and the Eth column as an example for explanation. After the replacement processing based on the granularity of row data, the order of row data in the generated first intermediate data is: the zeroth row, the third row, the first row and the second row. The connection paths of the multiple nodes corresponding to the zeroth row, the third row, the first row and the second row of the first intermediate data satisfy the shortest Hamiltonian path, and the amount of data of the first intermediate data after data compression processing is the smallest among all replacement schemes of the first data. It should be noted that each color block in Figure 5 represents data, such as a character string.
[0141] In one example, in conjunction with FIG5 , a replacement scheme for obtaining the first intermediate data is described. Please refer to FIG6 , which is a flow chart of a replacement processing method in an embodiment of the present application. The replacement processing method proposed in the embodiment of the present application includes:
[0142] G1. Map the row data of each row in the first data to a node to generate a node set.
[0143] For example, as shown in Figure 7, Figure 7 is a schematic diagram of mapping row data to nodes in an embodiment of the present application. The zeroth row of the first data is mapped to node zero, the first row of the first data is mapped to node one, the second row of the first data is mapped to node two, and the third row of the first data is mapped to node three. The above nodes zero, one, two, and three constitute a node set.
[0144] G2. Construct a directed complete graph based on the node set, where any two nodes in the node set in the directed complete graph are connected. The directed complete graph includes all nodes in the node set.
[0145] G3. Determine the weight of the edge connecting any two nodes in a directed complete graph.
[0146] For the specific method of determining the weight of the edge, please refer to the description of the above embodiment, which will not be repeated here. For example, the obtained directed complete graph with weights is shown in Figure 8, which is a schematic diagram of a directed complete graph in an embodiment of the present application.
[0147] G4. Initialize the feasible edge set and target edge set. The initialized feasible edge set includes all feasible edges in the directed complete graph, and the initialized target edge set is empty.
[0148] In step G4, an initialized feasible edge set is determined based on the directed complete graph. The initialized feasible edge set includes all feasible edges in the directed complete graph, and the feasible edges in the initialized feasible edge set are sorted by weight. In addition, an initialized target edge set is determined, and the initialized target edge set is empty.
[0149] G5. Randomly select one edge from the two edges with the smallest weight in the feasible edge set to generate the target edge set.
[0150] After step G5, execute step G6.
[0151] G6. Delete the edges of the target edge set from the feasible edge set.
[0152] After step G6, execute step G7.
[0153] G7. Check whether the subgraph consisting of the target edge set contains the shortest Hamiltonian path of the directed complete graph.
[0154] Repeat steps G5 to G6 and G7 until the subgraph consisting of the target edge set detected in step G7 includes the shortest Hamiltonian path of the directed complete graph. For example, as shown in Figure 9, Figure 9 is a schematic diagram of solving the shortest Hamiltonian path in an embodiment of the present application. A heuristic solution method is used for the directed complete graph shown in Figure 8 to solve the shortest Hamiltonian path of the directed complete graph. The specific solution method is as described in steps G3 to G7 above. The shortest Hamiltonian path of the directed complete graph shown in Figure 9 is: node zero → node three → node one → node two.
[0155] After it is determined that the subgraph composed of the target edge set includes the shortest Hamiltonian path of the directed complete graph, the process proceeds to step G8 to determine an index sequence based on the shortest Hamiltonian path of the directed complete graph.
[0156] G8. Determine an index sequence based on the shortest Hamiltonian path of the directed complete graph. The order of the index sequence is consistent with the order of the index values of the nodes in the shortest Hamiltonian path of the directed complete graph.
[0157] Exemplarily, the index sequence determined according to the shortest Hamiltonian path of the directed complete graph shown in FIG9 is: row 0→row 3→row 1→row 2 (or: zero, three, one, two).
[0158] G9. Perform permutation processing on the first data according to the order of the index sequence at the granularity of row data to generate first intermediate data. The order of the plurality of row data included in the first intermediate data is the same as the order of the index sequence.
[0159] Exemplarily, as shown in FIG5 , the order of the row data of the generated first intermediate data is consistent with the index sequence illustrated in FIG9 .
[0160] S3. Perform compression encoding processing on the first intermediate data to generate second data.
[0161] In step S3, the first intermediate data is compressed and encoded to generate second data. Specific methods for this compression and encoding include, but are not limited to, run-length coding (RLC), Huffman coding, differential coding, LZ77 coding, or Shannon-Fano coding. Taking run-length coding as an example, a string "AAAABBBCCDEEEE" consisting of 4 As, 3 Bs, 2 Cs, 1 D, and 4 Es can be compressed into 4A3B2C1D4E (compressing 14 characters into 10) through run-length coding.
[0162] In the embodiment of the present application, based on the replacement processing of row data, the long-tail distribution characteristics of the data in the table data can be effectively utilized to improve the compression rate of the data. The order of the multiple row data of the first intermediate data is consistent with the order indicated by the shortest Hamiltonian path connecting each node in the node set. A node in the node set is mapped by a row data of the first intermediate data. The multiple nodes included in the node set correspond to multiple row data of the first intermediate data, further highlighting the locality of the data, improving the compression rate of the data, and reducing the overhead of data movement during the replacement processing. The use of a greedy algorithm to solve the shortest Hamiltonian path of the node set can improve the encoding speed and the compression speed of data compression. This solution reduces the joint relocation amount of table data, improves the compression speed of table data, and improves the compression ratio performance of table data by exploiting the locality of table data.
[0163] In combination with the foregoing embodiments, the data compression method proposed in the embodiment of the present application can not only replace the table data according to the granularity of row data, but also divide and jointly process according to the granularity of column data, and use the correlation between the column data in the table data to reduce the redundancy between the column data in the table data, thereby improving the compression rate of the data. The specific method is introduced below, and the division and joint processing of the first intermediate data according to the granularity of column data is used as an example for explanation. It can be understood that the first data can also be divided and jointly processed according to the granularity of column data to generate the second intermediate data, and then the second intermediate data can be replaced according to the granularity of row data to generate the first intermediate data. The embodiment of the present application does not limit the execution order. Please refer to Figure 10, which is a flow chart of a division and joint processing method in the embodiment of the present application. The division and joint processing method proposed in the embodiment of the present application includes:
[0164] D1. Divide and jointly process the first intermediate data according to the granularity of column data to obtain second intermediate data.
[0165] Specifically, the first intermediate data is processed to obtain the second intermediate data, the second intermediate data including at least the first column data and the second column data, the first column data including one or more columns of column data of the first intermediate data, and the second column data including one or more columns of column data of the first intermediate data. For the sake of convenience, the above-mentioned processing of the first intermediate data can be referred to as partitioning and combining processing (or column combining processing, or column partitioning processing), which refers to dividing the multiple columns of column data of the table data into one or more columns of column data according to the granularity of the column data, and then merging the one or more columns of column data into one column of column data. For example, as shown in Tables 5 and 6, Table 5 is the first intermediate data, and the first intermediate data shown in Table 5 is partitioned and combined, and the generated second intermediate data is shown in Table 6.
[0166] Table 6
[0167] The following description is provided with reference to the first intermediate data shown in Table 4 and the second intermediate data shown in Table 5. The first intermediate data includes column data A, column data B, column data C, column data D, and column data E. The second intermediate data includes column data ABE and column data CD. Column data ABE is obtained by combining column data A, column data B, and column data D, while column data CD is obtained by combining column data C and column data E.
[0168] Optionally, the second intermediate data may further include second indication information, which is used to indicate the correspondence between the column data of the second intermediate data and the column data of the first intermediate data. For example, the second indication information is "A_B_E" and "C_D" in the first row of the data in Table 6. The second indication information "A_B_E" indicates that column data ABE is obtained by merging column data A, column data B, and column data E, and the second indication information "C_D" indicates that column data CD is obtained by merging column data C and column data D. It is understood that the second indication information can also be independent of the second intermediate data. For example, the first intermediate data can be divided and processed according to the granularity of the column data to output the second intermediate data and the second indication information.
[0169] There is often a strong correlation between the various column data of the table data. For example, if a column data in the table data is an IP address and another column data in the table data is a port number, then there is a strong correlation between the above two columns of data. In order to utilize the correlation between the column data of the table data, the first intermediate data is divided and jointly processed according to the granularity of the column data in step D1. The division and joint processing refers to dividing the multiple columns of data of the original table data into one or more column data sets, each column data set includes one or more column data, and then combining the one or more column data included in the column data set into one column data, so that the number of column data in the new table data obtained after the division and joint processing is less than the number of column data in the original table data. Taking the division and joint processing of column data A and column data B as an example, the gain of the division and joint processing of the column data of the table data can be expressed by the following formula: G(A,B)=(H(A)+H(B))-H(A,B)-δ;
[0170] Among them, G(A,B) is the gain of dividing and jointly processing column data A and column data B; H(A,B) is the joint entropy of column data AB after column data A and column data B are jointly (or merged) processed, and column data AB includes column data A and column data B; H(A) is the entropy of column data A, and H(B) is the entropy of column data B; δ is the computational cost of dividing and jointly processing column data A and column data B.
[0171] The following describes the algorithm part of the partition joint processing:
[0172] Taking the partitioning and joint processing of the first intermediate data as an example, the goal of partitioning and joint processing the first intermediate data according to the granularity of column data is: in a certain partitioning and joint scheme (referred to as the partitioning and joint scheme), the table data is partitioned and processed according to the partitioning and joint scheme, and then the sets of one or more column data obtained by the partitioning and joint processing are respectively joint processed to obtain one or more columns of new column data, and the one or more columns of new column data together constitute new table data (the table data obtained after the partitioning and joint processing is referred to as the second intermediate data). The compressed file obtained after the data compression processing of the second intermediate data has the best compression ratio compared to the compressed file obtained by any other partitioning and joint scheme (or the file size obtained after data compression is the smallest, or the data volume of the compressed file after data compression is the smallest). The optimized partitioning and joint scheme (or the mathematical optimization model of the partitioning and joint scheme, or the optimal solution of the partitioning and joint scheme, or the optimal solution of column partitioning, or the optimal solution of column joint) can be expressed as follows:
[0173] Wherein, M = {1, 2, ..., m} represents the set of all column data in the table data; represents all possibilities of partitioning the joint scheme, i.e., the solution space of the above optimization model problem; p is a scheme for performing partitioning and joint processing on the first intermediate data, and p includes one or more column data in the first intermediate data; It is a function of the value of p, which indicates the size of the compressed file obtained by performing data compression on the table data after partitioning and combining the table data according to the partitioning and combining scheme p. Represents the p[1]th column of the table data after the partitioning and combining process. The p[1]th column of the table data includes one or more column data in the first intermediate data indicated by the p[1]th partitioning and combining scheme. In particular, if |p| = 1, then the value function corresponds to the entropy of the column data, that is, v(p) = H(p); if |p| = 2, then the value function corresponds to two columns of data (for example: and ), namely:
[0174] For the above optimization model, there exists an optimal solution P * ={p1,p2,…,p j ,p j+1 …,p k} is ∪ i=1:k p i The optimal partitioning joint solution in the column joint, then for any j, there exists {p1,…,p j} is its sub-problem ∪ i=1:j p i The optimal partitioning joint solution in column joint, {p j+1 ,…,p k} is its subproblem ∪ i=j+1:k p i The optimal partitioning joint solution in column joint where k is a positive integer greater than or equal to 1 and less than or equal to m, ∪ i=1:k p i Column union refers to dividing and unioning the column data of the first column to the kth column in the first intermediate data, ∪ i=1:j p i Column union refers to dividing and unioning the column data of the first column to the jth column in the first intermediate data, ∪ i=j+1:k p i Column join refers to partitioning and joining column data from columns j+1 to k in the first intermediate data, where j is less than k. Column join refers to a scheme for partitioning and joining column data in at least two columns included in the first intermediate data. Column join includes one or more schemes for partitioning and joining column data in the first intermediate data.
[0175] make is the optimization function for p, and f(p) means: the size of the compressed file obtained by compressing the second intermediate data obtained by first partitioning the first intermediate data according to the partitioning scheme p and then performing the combined processing. Furthermore, the following recursive expression for f(p) can be obtained:
[0176] in, represents the possibility of partitioning the joint solution, belong f(p′) is the size of the first compressed data, which is the compressed data obtained by dividing and combining the first intermediate data according to the first column union and then performing data compression processing on the tabular data; f(p″) is the size of the second compressed data, which is the compressed data obtained by dividing and combining the first intermediate data according to the second column union and then performing data compression processing on the tabular data; the intersection between the column data in the first intermediate data corresponding to the first column union and the column data of the first intermediate data corresponding to the second column union is empty. For example, the first column union corresponds to the column data of the 1st to jth columns in the first intermediate data, and the second column union corresponds to the column data of the j+1th to kth columns in the first intermediate data.
[0177] According to the dynamic programming algorithm, the optimal solution of f(p) is obtained, namely P * , the optimal solution satisfies: the sum of the size of the first compressed data and the size of the second compressed data is the minimum value of the data after the first intermediate data is processed by arbitrary division and joint processing and data compression. * At least includes: a first column union and a second column union. After determining the first column union and the second column union in the optimal solution, first column data and second column data of the second intermediate data are determined according to the optimal solution, and finally second intermediate data is determined according to the first column data and the second column data. The first column data and the second column data constitute the second intermediate data, wherein the first column data corresponds to the partitioning and union scheme included in the first column union, the second column data corresponds to the partitioning and union scheme included in the second column union, the first column data includes one or more columns of column data, and the second column data includes one or more columns of column data.
[0178] It is understandable that the optimal solution may include more column join schemes, such as a third column join, a fourth column join, and a fifth column join, etc. Each column join included in the optimal solution includes one or more schemes for partitioning and joining column data. Based on the column join, it can be determined that one or more columns of column data of the first intermediate data are partitioned and joined to obtain one or more columns of new column data. The new column data obtained through the partitioning and joining includes the column data of one or more columns in the first intermediate data.
[0179] Below, taking the first intermediate data shown in Figure 5 as an example, the process of dividing and jointly processing the first intermediate data according to the granularity of column data to obtain the second intermediate data is explained. Please refer to Figure 11, which is a schematic diagram of the division and joint processing in an embodiment of the present application. The first intermediate data includes column data of column A, column B, column C, column D and column E. The column data of the above 5 columns are mapped to a weighted complete graph, and the weighted complete graph includes 5 nodes: node A, node B, node C, node D and node E, wherein column A is mapped to node A, column B is mapped to node B, column C is mapped to node C, column D is mapped to node D, and column E is mapped to node E. Any two nodes in the weighted complete graph are connected, and the weight between any two nodes is the gain of the two columns of column data corresponding to the two nodes for division and joint processing. For example, the weight between node A and node B is the gain of the column data of column A and the column data of column B for division and joint processing.
[0180] A dynamic programming algorithm is then used to find the optimal solution for the weighted complete graph. One possible solution is as follows: 1. The first intermediate data is partitioned and combined at a granularity of one column to obtain a set of column data: {A}, {B}, {C}, {D}, and {E}. For example, set {A} includes column data A. These sets of column data are then merged and compressed to obtain the compressed data size of each set.
[0181] 2. Divide the first intermediate data into two columns and perform joint processing to obtain a set of column data: {A, B}, {A, C}, {A, D}, {A, E}, {B, C}, ..., {D, E}, and so on. For example, the set {A, B} includes column data A and column data B. Then, merge these sets of column data to obtain column data AB, column data AC, column data AE, ..., column data DE, and so on. For example, column data AB includes column data A and column data B. Compression is performed to obtain the compressed data size of each set. The merged column data is then compressed to obtain the compressed data size of each set.
[0182] 3. Divide and jointly process the first intermediate data using a granularity of three columns. The optimal solution for partitioning and jointly processing the first intermediate data using a granularity of three columns can be solved by dividing it into two subproblems: Subproblem 1: Divide and jointly process the first intermediate data using a granularity of one column; Subproblem 2: Divide and jointly process the first intermediate data using a granularity of two columns. Using the above method, the size of the compressed data after compression is determined for each solution for partitioning and jointly processing the first intermediate data using a granularity of three columns.
[0183] 4. Divide the first intermediate data into four columns and perform joint processing. The optimal solution to dividing the first intermediate data into four columns and performing joint processing can be divided into two sub-problems for solution: Sub-problem 1 and Sub-problem 3. Divide the first intermediate data into three columns and perform joint processing, and Sub-problem 3 can be divided into two finer sub-problems for solution: Sub-problem 1 and Sub-problem 2. Alternatively, the optimal solution to dividing the first intermediate data into four columns and performing joint processing can be divided into two sub-problems 2. Through the above method, the size of the compressed data after compression processing is determined for each scheme for dividing the first intermediate data into four columns and performing joint processing.
[0184] 5. Divide the first intermediate data into 5 columns and perform joint processing. The optimal solution problem for dividing the first intermediate data into 5 columns and performing joint processing can be divided into two sub-problems for solution: Sub-problem 1 and Sub-problem 4. Divide the first intermediate data into 4 columns and perform joint processing. For Sub-problem 4, it can be divided into two sub-problems: Sub-problem 1 and Sub-problem 3, or two Sub-problems 2 for solution. Similarly, Sub-problem 3 can be divided into Sub-problem 1 and Sub-problem 2 for solution. Alternatively, the optimal solution problem for dividing the first intermediate data into 5 columns and performing joint processing can be divided into Sub-problems 2 and Sub-problem 3 for solution.
[0185] Using the above method, the size of the compressed data after compression processing for each scheme in which the first intermediate data is divided and processed at a column data granularity is determined. Then, from the multiple schemes, the scheme with the smallest sum of compressed data is selected as the optimal solution. For example, the optimal solution includes a first column join and a second column join, where the first column join indicates that columns A, B, and E are to be processed jointly, and the second column join indicates that columns C and D are to be processed jointly.
[0186] Based on the first column union, it is determined that the first column of data includes columns A, B, and E, which are the column data of columns A, B, and E combined into one column. Based on the second column union, it is determined that the second column of data includes columns C and D, which are the column data of columns C and D combined into one column. The first column of data and the second column of data together constitute the second intermediate data.
[0187] It should be noted that, to further improve the compression rate, the optimal solution for dividing and jointly processing the first intermediate data may be an optimal solution generated by a data compression device running the data compression method based on one or more offline first intermediate data. When the first data is compressed and encoded online, the first intermediate data (or first data) generated from the first data may be directly divided and jointly processed using the optimal solution generated offline to generate second intermediate data.
[0188] D2. Perform compression encoding processing on the second intermediate data to generate second data.
[0189] In step D2, the second intermediate data is compressed and coded to generate second data. Specific methods of the compression coding include, but are not limited to, run-length coding, Huffman coding, differential coding, LZ77 coding, or Shannon-Fanno coding.
[0190] It should be noted that in another possible implementation, after obtaining the first data, the first data can be divided and jointly processed according to the granularity of column data to obtain fifth intermediate data, where the fifth intermediate data includes at least third column data and fourth column data, where the third column data includes one or more columns of column data of the first data, and the fourth column data includes one or more columns of column data of the first data. The specific method for dividing and jointly processing the first data is similar to the method for dividing and jointly processing the first intermediate data, and is not described in detail here. Then, the fifth intermediate data is permuted according to the granularity of row data to generate sixth intermediate data. The specific method for permuting the fifth intermediate data according to the granularity of row data is similar to the method for permuting the first data according to the granularity of row data, and is not described in detail here.
[0191] In the embodiments of the present application, the column-based partitioning and joining processing can effectively utilize the correlation between column data, avoid column data redundancy in table data, and improve data compression rate. The dynamic programming algorithm is used to solve the column data partitioning and joining scheme, which can improve the encoding speed and the compression speed of data compression.
[0192] In combination with the foregoing embodiments, the data compression method proposed in the embodiments of the present application can also use a public dictionary to encode the tabular data, thereby saving dictionary storage overhead and improving the data compression rate. For different data files to be encoded (or referred to as encoding objects), the embodiments of the present application can construct a variety of different public dictionaries. In one example, the second intermediate data is encoded using a first public dictionary, and the first public dictionary is constructed from multiple second intermediate data obtained offline. In another example, the first data is encoded using a second public dictionary, and the second public dictionary is constructed from multiple first data obtained offline. In another example, the first intermediate data is encoded using a third public dictionary, and the third public dictionary is constructed from multiple first intermediate data obtained offline.
[0193] It should be noted that the data compression method of the embodiment of the present application is not limited in the order in which the aforementioned permutation processing based on the granularity of row data, the partitioning and joint processing based on the granularity of column data, and the encoding processing using a common dictionary are executed. The data compression method of the embodiment of the present application can also adopt any one or more of the aforementioned processing methods according to actual needs.
[0194] For example: the first data is encoded using the second public dictionary to generate fourth intermediate data, and then the fourth intermediate data is replaced according to the granularity of row data to generate first intermediate data, and finally the first intermediate data is compressed and encoded to output second data.
[0195] For another example: the first data is replaced to generate the first intermediate data, the first intermediate data is divided and combined to generate the second intermediate data, the second intermediate data is encoded using the first common dictionary to generate the third intermediate data, and finally the third intermediate data is compressed and encoded to output the second data.
[0196] For another example: the first data is replaced to generate the first intermediate data, the first intermediate data is divided and combined to generate the second intermediate data, the second intermediate data is encoded using the first common dictionary to generate the third intermediate data, and finally the third intermediate data is compressed and encoded to output the second data.
[0197] For another example, the first data is divided and combined to generate fifth intermediate data, where the fifth intermediate data includes at least third and fourth columns of data, where the third column of data includes one or more columns of data from the first data, and the fourth column of data includes one or more columns of data from the first data. Then, the fifth intermediate data is permuted to generate sixth intermediate data, where the order of the multiple rows of data included in the sixth intermediate data is different from the order of the multiple rows of data included in the first intermediate data. Finally, the sixth intermediate data is compressed and encoded to output the second data.
[0198] The following describes the specific method of encoding using a public dictionary and the specific method of constructing a public dictionary, taking the first public dictionary as an example. It should be noted that the specific method of encoding and constructing the second and third public dictionaries are similar to those of the first public dictionary and will not be described in detail here. Please refer to Figure 12, which is a flow chart of a public dictionary-based encoding method in an embodiment of the present application. The public dictionary-based encoding method proposed in this embodiment of the present application includes:
[0199] F1. Use the first public dictionary to encode the second intermediate data to obtain third intermediate data.
[0200] Since some columns of data in the tabular data may occupy a large storage space, a dictionary-based variable-length encoding can be used for compact representation. However, building a dictionary for each data file to be encoded results in a large storage overhead. Therefore, an embodiment of the present application proposes a public dictionary, which reduces the storage overhead of the dictionary by building one or more public dictionaries based on multiple data files to be encoded. The multiple data files to be encoded can be obtained offline, that is, the data files to be encoded are obtained from the historical files that have been encoded as the basis for building the public dictionary.
[0201] The following describes a specific method for encoding the second intermediate data using the first common dictionary to obtain the third intermediate data. For example, please refer to Figure 13, which is a schematic diagram of an encoding method based on a common dictionary in an embodiment of the present application. The data to be matched in the first common dictionary include: (0,0), (1,1), (0,1) and (1,0), wherein the encoding value (key) corresponding to the data to be matched (0,0) is 0, the encoding value (key) corresponding to the data to be matched (1,0) is 0, the encoding value (key) corresponding to the data to be matched (0,1) is 1, and the encoding value (key) corresponding to the data to be matched (1,1) is 1. The encoding process for the data to be matched to generate the encoding value is, and the encoding process includes but is not limited to: Huffman coding, entropy coding or arithmetic coding, etc. In the third intermediate data generated by encoding the second intermediate data using the first common dictionary, the 0th row of the first column of data is encoded as 0, the 0th row of the second column of data is encoded as 0, the 1st row of the first column of data is encoded as 1, the 1st row of the first column of data is encoded as 0, and so on.
[0202] F2. Perform compression encoding processing on the third intermediate data.
[0203] For example, please refer to Figure 14, which is a schematic diagram of compression encoding processing of the third intermediate data in an embodiment of the present application. In step F2, compression encoding processing is performed on the third intermediate data to generate second data. Specific methods for compression encoding processing include, but are not limited to, run-length encoding, Huffman encoding, differential encoding, LZ77 encoding, or Shannon-Fano encoding.
[0204] In the embodiment of the present application, variable length coding is performed online based on the constructed public dictionary to achieve a compact representation of the data and improve the data compression rate. In addition, the storage overhead of the dictionary can also be reduced.
[0205] In conjunction with the above embodiments, the following describes the process of constructing a public dictionary. Taking the construction of the first public dictionary as an example, it is understood that the method of constructing the second public dictionary or the third public dictionary is similar to that of constructing the first public dictionary, and will not be described in detail here. Please refer to Figure 15, which is a flow chart of a method for constructing the first public dictionary in an embodiment of the present application. A method for constructing the first public dictionary in an embodiment of the present application includes:
[0206] H1. Obtain multiple second intermediate data.
[0207] For example, as shown in Figure 16a, which is a schematic diagram of constructing the first public dictionary in an embodiment of the present application, a plurality of second intermediate data are obtained from the encoded historical files in an offline manner, and then the plurality of second intermediate data are used as the basis for constructing the first public dictionary.
[0208] H2. Generate a dictionary table based on the combination of the plurality of second intermediate data, where the dictionary table includes data of the plurality of second intermediate data.
[0209] H3. Divide the multiple to-be-matched data included in the dictionary table into one or more to-be-matched data sets according to the statistical data amount of the to-be-matched data in the dictionary table.
[0210] For example, as shown in Figure 17, Figure 17 is a diagram illustrating data frequency statistics for a dictionary table in an embodiment of the present application. The occurrence frequency of each to-be-matched data item in the dictionary table is counted, and the statistical quantity of the to-be-matched data item in the dictionary table is determined. The statistical quantity indicates the total amount of to-be-matched data item in the dictionary table, or in other words, the total number of to-be-matched data items in the dictionary table.
[0211] H4. Encode the data to be matched and generate a code value corresponding to the data to be matched.
[0212] The encoding process in step H4 includes but is not limited to: Huffman coding, entropy coding or arithmetic coding.
[0213] H5. Construct one or more first common dictionaries based on the data to be matched and the code values corresponding to the data to be matched, where each first common dictionary includes a set of data to be matched.
[0214] In step H5, the plurality of to-be-matched data included in the dictionary table are divided into one or more to-be-matched data sets according to the statistical data amount of the to-be-matched data, and each to-be-matched data set includes at least one to-be-matched data.
[0215] For example, the dictionary table includes a total of 300 data to be matched, which are sorted according to the frequency of occurrence of each data to be matched. The 1st to 100th data to be matched are divided into data set #1 to be matched; the 101st to 200th data to be matched are divided into data set #2 to be matched; and the 201st to 300th data to be matched are divided into data set #3 to be matched. Then, based on the data set #1 to be matched and the corresponding coding value, the first public dictionary #1 is constructed; the data set #2 to be matched and the corresponding coding value, the first public dictionary #2 is constructed; the data set #3 to be matched and the corresponding coding value, the first public dictionary #3 is constructed, and finally three first public dictionaries (first public dictionary #1, first public dictionary #2 and first public dictionary #3) are output. The above method reduces the overhead of querying dictionaries, saves encoding time, and improves the efficiency of compressing data.
[0216] Optionally, the storage device may also continuously acquire multiple second intermediate data during operation and update the first public dictionary based on changes in data distribution in the multiple second intermediate data, so that the first public dictionary can adapt to data changes and improve data compression rate.
[0217] In one possible implementation, the encoding device (or decoding device) runs on a local computing device (or storage device), the local computing device reports the second intermediate data to the cloud, the cloud constructs a first public dictionary based on the multiple second intermediate data reported by one or more local computing devices, and then the cloud sends the first public dictionary to the local computing device.
[0218] For example, as shown in FIG16b , FIG16b is a schematic diagram of a scenario for constructing a public dictionary in an embodiment of the present application. Cloud computing is a service related to information technology, software, and the Internet. Cloud computing brings together multiple computing resources to form a computing resource sharing pool, which is also called the "cloud" (or cloud). Through software, automated management is achieved, and tenants can obtain resources on the "cloud" at any time according to demand. In theory, the resources on the "cloud" can be expanded infinitely. In FIG16b , the cloud is deployed in a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. The local computing device illustrated in FIG16b is other computing devices compared to the cloud. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.
[0219] In Figure 16b, the cloud establishes a communication connection with one or more local computing devices, and the local computing devices generate one or more second intermediate data during the execution of the data compression method proposed in the embodiment of the present application. One or more local computing devices report multiple second intermediate data to the cloud, and then the cloud generates a first public dictionary based on the multiple second intermediate data. The cloud can also continuously update the first public dictionary based on the second intermediate data reported by the local computing devices. The cloud can send the first public dictionary or the updated first public dictionary to the local computing device for the local computing device to perform data compression or decompression based on the first public dictionary.
[0220] In another possible implementation, the encoding device (or decoding device) runs in the cloud. During execution of the data compression method proposed in the embodiment of the present application, the cloud generates a plurality of second intermediate data, and then generates a first common dictionary based on the plurality of second intermediate data. The cloud may also continuously update the first common dictionary based on the generated second intermediate data. The cloud performs data compression or decompression based on the generated first common dictionary.
[0221] In another possible implementation, the encoding device (or decoding device) runs on a local computing device. In the process of executing the data compression method proposed in the embodiment of the present application, the local computing device generates multiple second intermediate data, and then generates a first public dictionary based on the multiple second intermediate data. The local computing device can also continuously update the first public dictionary based on the generated second intermediate data. The local computing device performs data compression or decompression based on the generated first public dictionary. The local computing device can also report the first public dictionary to the cloud, so that the cloud distributes the first public dictionary to other local computing devices for use.
[0222] The data compression method proposed in the embodiment of the present application is generally applied to a storage device having an encoding module. In a possible application scenario, the storage device having an encoding module also has a decoding module. The storage device can execute the data compression method proposed in the embodiment of the present application and the data decompression method corresponding to the data compression method. The present application does not limit the storage device that executes the above-mentioned data compression method or the above-mentioned data decompression method. The data compression method proposed in the embodiment of the present application is disclosed only as a data compression method in this application to avoid redundancy. It should be understood that the above-mentioned data compression method can also be parsed as a data decompression method, that is, the storage device having a decoding module executes in reverse according to the method of the storage device having an encoding module, and its technical means are similar to those of the storage device with the encoding module. Then the data decompression method executed corresponding to the data compression method also falls within the scope of protection of the present application.
[0223] In a possible example, please refer to Figure 18, which is a flow chart of an embodiment of a data decompression method in an embodiment of the present application. It should be noted that the data decompression method in the embodiment of the present application is the reverse execution of the aforementioned data compression method. The data decompression method illustrated in Figure 18 is the reverse execution of the following data compression method: 1. Processing the first data to generate first intermediate data; 2. Processing the first intermediate data to generate second intermediate data; 3. Using the first public dictionary to encode the second intermediate data to generate third intermediate data; 4. Compressing and encoding the third intermediate data to generate second data. It can be understood that for different possible implementations of the data compression method in the embodiment of the present application, there are also different possible implementations of the corresponding data decompression method, which is not limited by the embodiment of the present application.
[0224] Taking the decompression device executing the data decompression method as an example, FIG18 shows a data decompression method, including:
[0225] R1. The decompression device obtains the second data.
[0226] R2. The decompression device decompresses the second data to generate third intermediate data.
[0227] It should be noted that the specific method of the decompression process in the embodiments of the present application is the reverse execution of the compression encoding process. For example, the decompression process is the reverse execution of the aforementioned run-length encoding, Huffman encoding, differential encoding, LZ77 encoding, or Shannon-Fanno encoding.
[0228] R3. The decompression device uses the first public dictionary to decode the third intermediate data to generate second intermediate data.
[0229] R4. The decompression device processes the second intermediate data to generate first intermediate data.
[0230] In one possible implementation, the decompression device obtains second indication information. Then, based on the second indication information, it determines the column data of the first intermediate data corresponding to the column data of the second intermediate data. Based on the second indication information, the second intermediate data is split to generate the first intermediate data. Exemplarily, the second intermediate data is as shown in Table 6, and the decompression device obtains the second indication information "A_B_E" and "C_D" from the second intermediate data. Based on the second indication information "A_B_E", it is determined that the column data ABE of the second intermediate data corresponds to the column data A, column data B, and column data E of the first intermediate data. Then, the column data ABE is split into three columns of column data (column data A, column data B, and column data E). Similarly, based on the second indication information "C_D", it is determined that the column data CD of the second intermediate data corresponds to the column data C and column data D of the first intermediate data. Then, the column data CD is split into two columns of column data (column data C and column data D). Finally, the first intermediate data is output, and the first intermediate data includes: column data A, column data B, column data C, column data D, and column data E.
[0231] R5. Process the first intermediate data to generate first data.
[0232] In one possible implementation, the decompression device obtains first indication information. Then, based on the first indication information, the order of each row of row data in the first intermediate data in the first data is determined. Then, based on the first indication information, the first intermediate data is subjected to row permutation processing to restore and generate the first data. Exemplarily, the first intermediate data is shown in Table 5. The decompression device obtains the first indication information "{3,2,1,4}" from the first intermediate data, that is, the "serial number" in the first column of Table 5. Then, based on the first indication information, the row data of the first intermediate data is permuted so that the "serial number" in the first column after the permutation processing is {1,2,3,4}. The tabular data output after the permutation processing is used as the first data. For example, the first intermediate data shown in Table 5 is permuted to generate the first data shown in Table 4.
[0233] In the embodiment of the present application, the above-mentioned data decompression method can improve the decoding speed and the data decompression speed.
[0234] In combination with the foregoing embodiments, an application scenario proposed in an embodiment of the present application is described below. Please refer to Figure 19a, which is a schematic diagram of an application scenario in an embodiment of the present application. Figure 19a illustrates the specific process of data compression performed by the data compression method proposed in an embodiment of the present application. First, the original "parquet" file is read. "Parquet" is a columnar storage file format that is widely used in the field of data storage. Then, the original parquet file is encoded based on the public dictionary using a public dictionary table to generate the fourth intermediate data (the lower left table data in Figure 19a). Then, the fourth intermediate data is partitioned and combined according to the optimal solution of the partition-join scheme to obtain the second intermediate data (the lower right table data in Figure 19a). Then, the second intermediate data is permuted according to the granularity of the row data to generate the first intermediate data (the upper right table data in Figure 19a). Finally, the first intermediate data is run-length encoded to obtain the compressed parquet file.
[0235] Exemplarily, please refer to Figure 19b, which is a schematic diagram of an application scenario in an embodiment of the present application. Figure 19b illustrates the specific process of data decompression by the data decompression method proposed in an embodiment of the present application. First, the compressed parquet file is read. Then, the compressed parquet file is subjected to reverse decoding processing using run-length encoding to restore and generate the first intermediate data (the upper right table data in Figure 19b). Then, the first intermediate data is restored in the reverse order of the permutation processing according to the granularity of the row data to restore and generate the second intermediate data (the lower right table data in Figure 19b). Then, the second intermediate data is divided and combined according to the granularity of the column data to restore and generate the third intermediate data (the lower left table data in Figure 19b). Finally, the third intermediate data is decoded using a public dictionary to restore the original parquet file (the upper left table data in Figure 19b).
[0236] On the basis of the embodiments corresponding to FIG. 3 to FIG. 19 b , in order to better implement the above-mentioned solutions of the embodiments of the present application, relevant equipment for implementing the above-mentioned solutions is also provided below.
[0237] Please refer to Figure 20, which is a schematic diagram of the structure of a data compression device provided in an embodiment of the present application. The data compression device includes:
[0238] The transceiver unit 2002 is configured to obtain first data, where the first data is table data consisting of a plurality of rows of row data and a plurality of columns of column data;
[0239] The processing unit 2001 is configured to process the first data to generate first intermediate data, wherein the order of the plurality of rows of data included in the first intermediate data is different from the order of the plurality of rows of data included in the first data.
[0240] The processing unit 2001 is further configured to perform compression encoding processing on the first intermediate data to generate second data.
[0241] In one possible implementation,
[0242] The processing unit 2001 is further configured to map row data of each row in the first data to a node to generate a node set, where the node set includes a plurality of nodes;
[0243] The processing unit 2001 is also used to process the first data according to the connection order of the multiple nodes included in the node set to generate the first intermediate data, and the order of the multiple row data of the first intermediate data is the same as the connection order of the multiple nodes included in the node set.
[0244] In a possible implementation, the connection order of the multiple nodes included in the node set is the same as the order indicated by the shortest Hamiltonian path connecting the nodes in the node set.
[0245] In one possible implementation,
[0246] The first intermediate data further includes first indication information, where the first indication information is used to indicate the order of each row of data in the first intermediate data in the first data.
[0247] In one possible implementation,
[0248] The processing unit 2001 is further configured to process the first intermediate data to obtain second intermediate data, where the second intermediate data includes at least a first column of data and a second column of data, where the first column of data includes one or more columns of column data of the first intermediate data, and the second column of data includes one or more columns of column data of the first intermediate data;
[0249] The processing unit 2001 is further configured to perform compression encoding processing on the second intermediate data to generate the second data.
[0250] In a possible implementation, the second intermediate data further includes second indication information, where the second indication information is used to indicate a correspondence between column data of the second intermediate data and column data of the first intermediate data.
[0251] In one possible implementation,
[0252] The processing unit 2001 is further configured to encode the second intermediate data using a first common dictionary to obtain third intermediate data, wherein the first common dictionary is constructed from a plurality of the second intermediate data, and the plurality of the second intermediate data corresponds to a plurality of the first data;
[0253] The processing unit 2001 is further configured to perform compression encoding processing on the third intermediate data to generate the second data.
[0254] In one possible implementation,
[0255] The processing unit 2001 is further configured to encode the first data using a second common dictionary to obtain fourth intermediate data, where the second common dictionary is constructed from a plurality of the first data;
[0256] The processing unit 2001 is further configured to process the fourth intermediate data to generate the first intermediate data, where the order of the multiple rows of data included in the first intermediate data is different from the order of the multiple rows of data included in the fourth intermediate data.
[0257] In one possible implementation,
[0258] The processing unit 2001 is further configured to construct a directed complete graph based on the node set, wherein any two nodes in the node set in the directed complete graph are connected by an edge;
[0259] The processing unit 2001 is further configured to solve the shortest Hamiltonian path of the directed complete graph;
[0260] The processing unit 2001 is further configured to determine an index sequence based on the shortest Hamiltonian path of the directed complete graph, wherein the index sequence includes a series of ordered index values, each index value in the index sequence corresponds to a node in the node set, and an order of the index sequence is consistent with an order of the index values of the nodes in the shortest Hamiltonian path of the directed complete graph;
[0261] The processing unit 2001 is further configured to process the first data according to the index sequence to generate the first intermediate data, wherein the order of the plurality of row data included in the first intermediate data is the same as the order of the index sequence.
[0262] In one possible implementation,
[0263] The transceiver unit 2002 is configured to obtain a plurality of the second intermediate data;
[0264] The processing unit 2001 is also used to construct the first common dictionary based on multiple second intermediate data, the first common dictionary including the data to be matched and the coding values corresponding to the data to be matched, and the data to be matched in the first common dictionary corresponds to the column data of any one or more columns in the second intermediate data.
[0265] In one possible implementation,
[0266] The transceiver unit 2002 is further configured to obtain a plurality of the second intermediate data;
[0267] The processing unit 2001 is further configured to update the first public dictionary according to distribution changes of the data in the plurality of second intermediate data.
[0268] In one possible implementation,
[0269] The transceiver unit 2002 is further configured to obtain a plurality of the first data;
[0270] The processing unit 2001 is also used to construct the second common dictionary based on multiple first data, the second common dictionary including the data to be matched and the coding values corresponding to the data to be matched, and the data to be matched in the second common dictionary corresponds to the column data of any one or more columns in the first data.
[0271] In one possible implementation,
[0272] The transceiver unit 2002 is further configured to obtain a plurality of the first data;
[0273] The processing unit 2001 is further configured to update the second public dictionary according to distribution changes of the data in the plurality of first data.
[0274] In one possible implementation,
[0275] The processing unit 2001 is further configured to perform run-length encoding on the first intermediate data to generate the second data.
[0276] Please refer to Figure 21, which is a schematic diagram of the structure of a decompression device provided in an embodiment of the present application. The decompression device includes:
[0277] The transceiver unit 2102 is configured to obtain second data, where the second data includes table data consisting of a plurality of rows of row data and a plurality of columns of column data;
[0278] The processing unit 2101 is configured to decompress the second data to generate first intermediate data;
[0279] The processing unit 2101 is further configured to process the first intermediate data to generate first data, where the order of the multiple rows of data included in the first intermediate data is different from the order of the multiple rows of data included in the first data.
[0280] In one possible implementation,
[0281] The transceiver unit 2102 is further configured to obtain first indication information from the first intermediate data, where the first indication information is used to indicate the order of each row of data in the first intermediate data in the first data;
[0282] The processing unit 2101 is further configured to process the first intermediate data according to the first indication information to generate the first data.
[0283] In one possible implementation,
[0284] The processing unit 2101 is further configured to decompress the second data to generate second intermediate data, where the second intermediate data includes at least a first column of data and a second column of data;
[0285] The processing unit 2101 is further used to process the second intermediate data to generate the first intermediate data, wherein one or more columns of the first intermediate data are obtained from the first column data, and one or more columns of the first intermediate data are obtained from the second column data.
[0286] In one possible implementation,
[0287] The transceiver unit 2102 is further configured to obtain second indication information from the second intermediate data, where the second indication information is used to indicate a correspondence between column data of the second intermediate data and column data of the first intermediate data;
[0288] The processing unit 2101 is further configured to process the second intermediate data according to the second indication information to generate the first intermediate data.
[0289] In one possible implementation,
[0290] The processing unit 2101 is further configured to decompress the second data to generate third intermediate data;
[0291] The processing unit 2101 is further configured to use a first common dictionary to decode the third intermediate data to obtain the second intermediate data, where the first common dictionary is constructed from a plurality of the second intermediate data, and the plurality of the second intermediate data corresponds to a plurality of the first data.
[0292] In one possible implementation,
[0293] The processing unit 2101 is further configured to process the first intermediate data to obtain fourth intermediate data, wherein the order of the plurality of row data included in the first intermediate data is different from the order of the plurality of row data included in the fourth intermediate data;
[0294] The processing unit 2101 is further configured to use a second common dictionary to decode the fourth intermediate data to obtain the first data, where the second common dictionary is constructed from a plurality of the first data.
[0295] In one possible implementation,
[0296] The processing unit 2101 is further configured to perform run-length encoding decoding processing on the second data to generate the first intermediate data.
[0297] Please refer to Figure 22, which is a schematic diagram of the structure of a storage device 2201 provided in an embodiment of the present application. The storage device 2201 can be the data compression device or decompression device in the aforementioned embodiment. As shown in Figure 22, the storage device 2201 includes a processor 2203, which is coupled to a system bus 2205. The processor 2203 can be one or more processors, each of which can include one or more processor cores. A display adapter (video adapter) 2207 can drive a display 2209, which is coupled to the system bus 2205. The system bus 2205 is coupled to an input / output (I / O) bus via a bus bridge 2211. An I / O interface 2226 is coupled to the I / O bus. The I / O interface 2226 communicates with various I / O devices, such as an input device 2217 (e.g., a touch screen), an external memory 2210 (e.g., a hard disk, floppy disk, optical disk, or USB flash drive), a multimedia interface, etc. A transceiver 2223 (which can send and / or receive radio communication signals) and an external USB port 2225. Optionally, the interface connected to the I / O interface 2226 can be a USB interface.
[0298] The processor 2203 may be any conventional processor, including a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, or a combination thereof. Alternatively, the processor may be a dedicated device such as an ASIC.
[0299] The hard drive interface 2231 is coupled to the system bus 2205. The hard drive interface is connected to the hard drive 2233. The internal memory 2235 is coupled to the system bus 2205. The data running in the internal memory 2235 may include the operating system (OS) 2237 of the storage device 2201, the application program 2243, and the scheduler.
[0300] The processor 2203 can communicate with the internal memory 2235 through the system bus 2205, and retrieve instructions and data in the application program 2243 from the internal memory 2235, thereby implementing program execution.
[0301] An operating system consists of a shell 2239 and a kernel 2241. Shell 2239 is an interface between the user and the operating system's kernel. The shell is the outermost layer of the operating system. The shell manages the interaction between the user and the operating system: it waits for user input, interprets user input to the operating system, and processes various operating system output.
[0302] The kernel 2241 consists of the parts of the operating system that manage memory, files, peripherals, and system resources. The kernel 2241 interacts directly with the hardware. The operating system kernel typically runs processes and provides communication between processes, CPU time slice management, interrupts, memory management, and I / O management.
[0303] An embodiment of the present application also provides a computer program product, which, when executed on a computer, enables the computer to execute the steps executed by the aforementioned storage device, or enables the computer to execute the steps executed by the aforementioned computing device.
[0304] A computer-readable storage medium is also provided in an embodiment of the present application. The computer-readable storage medium stores a program for signal processing. When the computer-readable storage medium is run on a computer, the computer executes the steps executed by the aforementioned storage device, or the computer executes the steps executed by the aforementioned storage device or computing device.
[0305] The storage device or computing device provided in the embodiments of the present application may specifically be a chip, and the chip includes: a processing unit and a communication unit, wherein the processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, a pin, or a circuit. The processing unit may execute computer-executable instructions stored in the storage unit so that the chip in the storage device executes the compilation method described in the above embodiment. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc. The storage unit may also be a storage unit located outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
[0306] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0307] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0308] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0309] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0310] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory, a random access memory, a magnetic disk or an optical disk.
Claims
1. A data compression method, characterized in that, Including: Obtain first data, where the first data is tabular data composed of row data in multiple rows and column data in multiple columns; Process the first data to generate first intermediate data, where the sorting of the multiple row data included in the first intermediate data is different from the sorting of the multiple row data included in the first data; Perform compression encoding processing on the first intermediate data to generate second data.
2. The method according to claim 1, characterized in that, Processing the first data to generate the first intermediate data includes: Map the row data of each row in the first data to a node to generate a node set, where the node set includes multiple nodes; Process the first data according to the connection order of the multiple nodes included in the node set to generate the first intermediate data, where the sorting of the multiple row data of the first intermediate data is the same as the connection order of the multiple nodes included in the node set.
3. The method according to claim 2, wherein The connection order of the multiple nodes included in the node set is the same as the order indicated by the shortest Hamiltonian path connecting each node in the node set.
4. The method according to any one of claims 1-3, characterized in that, The first intermediate data further includes first indication information, where the first indication information is used to indicate the sorting of each row of row data in the first intermediate data in the first data.
5. The method according to any one of claims 1-4, characterized in that, Performing compression encoding processing on the first intermediate data to generate the second data includes: Process the first intermediate data to obtain second intermediate data, where the second intermediate data includes at least first column data and second column data, the first column data includes one or more columns of the column data of the first intermediate data, and the second column data includes one or more columns of the column data of the first intermediate data; Perform compression encoding processing on the second intermediate data to generate the second data.
6. The method according to claim 5, wherein The second intermediate data further includes second indication information, where the second indication information is used to indicate the correspondence between the column data of the second intermediate data and the column data of the first intermediate data.
7. The method according to claim 5 or 6, characterized in that Performing compression encoding processing on the second intermediate data to generate the second data includes: Encode the second intermediate data using a first common dictionary to obtain third intermediate data, where the first common dictionary is constructed from multiple second intermediate data, and the multiple second intermediate data correspond to multiple first data; Perform compression encoding processing on the third intermediate data to generate the second data.
8. The method according to any one of claims 1-4, characterized in that, Performing permutation processing on the first data according to the granularity of row data to generate the first intermediate data includes: Encode the first data using a second common dictionary to obtain fourth intermediate data, where the second common dictionary is constructed from multiple first data; Process the fourth intermediate data to generate the first intermediate data, where the sorting of the multiple row data included in the first intermediate data is different from the sorting of the multiple row data included in the fourth intermediate data.
9. The method according to any one of claims 1-8, characterized in that, Processing the first data according to the connection order of the multiple nodes included in the node set to generate the first intermediate data includes: Construct a directed complete graph according to the node set, where any two nodes in the node set of the directed complete graph are connected by an edge; Solve the shortest Hamiltonian path of the directed complete graph; Determine an index sequence according to the shortest Hamiltonian path of the directed complete graph, where the index sequence includes a series of ordered index values, each of the index values in the index sequence corresponds to a node in the node set, and the order of the index sequence is consistent with the order of the index values of the nodes in the shortest Hamiltonian path of the directed complete graph; Process the first data according to the index sequence to generate the first intermediate data, and the sorting of the multiple row data included in the first intermediate data is the same as the sorting of the index sequence.
10. The method according to claim 7, wherein The method further includes: Obtain a plurality of the second intermediate data; Construct the first common dictionary according to the plurality of the second intermediate data, where the first common dictionary includes data to be matched and the encoding values corresponding to the data to be matched, and the data to be matched in the first common dictionary corresponds to the column data of any one column or multiple columns in the second intermediate data.
11. The method according to claim 10, wherein The method further includes: Obtain a plurality of the second intermediate data; Update the first common dictionary according to the distribution change of the data in the plurality of the second intermediate data.
12. The method according to claim 8, wherein The method further includes: Obtain a plurality of the first data; Construct the second common dictionary according to the plurality of the first data, where the second common dictionary includes data to be matched and the encoding values corresponding to the data to be matched, and the data to be matched in the second common dictionary corresponds to the column data of any one column or multiple columns in the first data.
13. The method according to claim 12, characterized in that, The method further includes: Obtain a plurality of the first data; Update the second common dictionary according to the distribution change of the data in the plurality of the first data.
14. The method according to any one of claims 1-4, characterized in that, Perform compression encoding processing on the first intermediate data to generate the second data, including: Perform run-length encoding processing on the first intermediate data to generate the second data.
15. A data decompression method, characterized in that, Includes: Obtain the second data, where the second data includes tabular data composed of multiple rows of row data and multiple columns of column data; Perform decompression processing on the second data to generate the first intermediate data; Process the first intermediate data to generate the first data, and the sorting of the multiple row data included in the first intermediate data is different from the sorting of the multiple row data included in the first data.
16. The method according to claim 15, characterized in that, Process the first intermediate data to generate the first data, including: Obtain first indication information from the first intermediate data, where the first indication information is used to indicate the sorting of each row of row data in the first intermediate data in the first data; Process the first intermediate data according to the first indication information to generate the first data.
17. According to the method of claim 15 or 16, perform decompression processing on the second data to generate the first intermediate data, including: Perform decompression processing on the second data to generate the second intermediate data, where the second intermediate data includes at least the first column data and the second column data; Process the second intermediate data to generate the first intermediate data, where one or more columns of column data of the first intermediate data are obtained from the first column data, and one or more columns of column data of the first intermediate data are obtained from the second column data.
18. The method according to claim 17, wherein Processing the second intermediate data to generate the first intermediate data includes: Obtain second indication information from the second intermediate data, where the second indication information is used to indicate the correspondence between the column data of the second intermediate data and the column data of the first intermediate data; Process the second intermediate data according to the second indication information to generate the first intermediate data.
19. The method according to claim 17 or 18, characterized in that, Decompressing the second data to generate the second intermediate data includes: Decompress the second data to generate third intermediate data; Decode the third intermediate data using a first common dictionary to obtain the second intermediate data, where the first common dictionary is constructed from multiple second intermediate data, and the multiple second intermediate data correspond to multiple first data.
20. The method according to claim 15 or 16, characterized in that Processing the first intermediate data to generate the first data includes; Process the first intermediate data to obtain fourth intermediate data, where the sorting of the multiple row data included in the first intermediate data is different from the sorting of the multiple row data included in the fourth intermediate data; Decode the fourth intermediate data using a second common dictionary to obtain the first data, where the second common dictionary is constructed from multiple first data.
21. The method according to any one of claims 15-20, characterized in that, Decompressing the second data to generate the first intermediate data includes: Perform decoding processing corresponding to run-length encoding on the second data to generate the first intermediate data.
22. An encoding device, characterized in that, The encoding device is used to execute the method according to any one of claims 1 to 14.
23. A decompression device, characterized in that, The encoding device is used to execute the method according to any one of claims 15 to 21.
24. A storage system, characterized in that, Including: The encoding device according to claim 22, and / or, the decompression device according to claim 23.
25. A computing system, characterized in that, The computing system includes an encoding device and a decompression device, where the encoding device is used to execute the method according to any one of claims 1 to 14, and the decompression device is used to execute the method according to any one of claims 15 to 21.
26. A chip, characterized in that, Including: An interface circuit and a processing circuit, the interface circuit and the processing circuit are connected, and the chip is used to execute the method according to any one of claims 1 to 14, and / or, execute the method according to any one of claims 15 to 21.
27. A computer program product, characterized in that, The computer program product stores instructions that, when executed by a computer, cause the computer to implement the method according to any one of claims 1 to 14, or the method according to any one of claims 15 to 21.
Citation Information
Patent Citations
Method for packing data, method for unpacking data, coder and decoder
CN103326732A
Executable code compression method of embedded type system and code uncompressing system
CN104331269A
Coding method and related equipment
CN112398484A
Data compression techniques
CN113688127A
Data processing method and device
CN115483935A