Direct reading of compressed data into in-memory repository
By directly loading compressed data into an in-memory repository and using transcoding and cardinality clustering techniques, the performance inefficiency during cold starts is solved, resulting in faster query response times and lower computational resource consumption.
Patent Information
- Application Number
- CN202480020007.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-05-31
- Filing Date
- 2024-04-29
- Publication Date
- 2025-11-11
AI Technical Summary
Existing cloud-based analytics query engines suffer from inefficiency and long cold start times due to the decompression and recompression of compressed data, failing to effectively utilize the performance advantages of in-memory repositories.
The compressed data is directly loaded into the in-memory repository. The file compression scheme is converted into an in-memory compression scheme by a transcoder, and cardinality clustering is used to accelerate the transcoding process, avoid decompression operations, and improve memory access efficiency.
It significantly reduces cold start time, lowers computational burden and power consumption, and improves memory access efficiency and query performance.
Smart Images

Figure CN120936995A_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 500,567, filed May 5, 2023, entitled “READING COMPRESSED DATA DIRECTLYINTO AN IN-MEMORY STORE”, the disclosure of which is incorporated herein by reference in its entirety. Background Technology
[0003] Common cloud-based architectures used for analytics query engines decouple compute resources from persistent storage in the data lake using open storage formats. Data is stored in compressed formats, causing performance issues during cold starts, where the data required to respond to a query must be loaded from storage into an initially empty in-memory repository (e.g., an in-memory column repository). For example, responding to a query from a cold start can take 40 times (or more) longer than responding to a query using a fully populated in-memory repository. Two approaches are typically used.
[0004] In one approach, compressed data is decompressed (e.g., decomposed) upon reading and loaded into an in-memory repository in its decompressed state. In this first approach, the in-memory repository consumes a significant amount of memory due to its nature of not supporting compressed data. Vectorized query execution within this in-memory repository exhibits reduced performance compared to utilizing a compressed in-memory repository.
[0005] Another approach uses an in-memory repository with a compressed data format, which offers higher memory efficiency and leverages the performance advantages of a vectorized query execution kernel. However, due to the difference between the compression schemes used in persistent storage and those used in the in-memory repository, the read process must decompress the data from the persistent storage format into a decompression buffer and recompress the decompressed data using a compression scheme compatible with the in-memory repository. This increases cold start time in addition to consuming memory for the decompression buffer. Both approaches are inefficient. Summary of the Invention
[0006] The disclosed examples are described in detail below with reference to the accompanying drawings. The following summary is provided to illustrate some of the examples disclosed herein. However, this is not to imply that all examples are limited to any particular configuration or sequence of operations.
[0007] An example solution for directly reading compressed data into an in-memory repository involves reading compressed data from a file, where the compressed data in the file has a stored compression scheme with a stored compression dictionary. The compressed data is then loaded into the in-memory repository without decompressing it. The compressed data in the in-memory repository has an in-memory compression scheme. A query is performed on the compressed data in the in-memory repository, and the query results are returned. In some examples, the in-memory compression scheme differs from the stored compression scheme, while in other examples, the in-memory compression scheme and the stored compression scheme are the same. In some examples, the in-memory compression scheme uses a stored compression dictionary, and in others, the in-memory compression scheme uses an in-memory compression dictionary. Attached Figure Description
[0008] The disclosed examples are described in detail below with reference to the accompanying drawings listed below:
[0009] Figure 1 The illustration shows an example architecture that advantageously provides direct reading of compressed data into an in-memory repository without decompression;
[0010] Figure 2 The illustration shows an example file for storing compressed data, such as... Figure 1 Used in the example architecture;
[0011] Figure 3A The diagram shows... Figure 2 Exemplary files and in-memory repositories (such as those that can be stored in memory) Figure 1 Comparisons between (used in the example architectures);
[0012] Figure 3B The diagram illustrates the target Figure 3A More details on the comparisons;
[0013] Figure 3C The diagram shows... Figure 2 Exemplary files and in-memory repositories (such as those that can be stored in memory) Figure 1 Another comparison between (those used in the example architectures);
[0014] Figure 4 The illustration shows an example of bit-packing compression, such as... Figure 1 Used in the example architecture;
[0015] Figure 5 The illustration shows an example of dictionary compression, such as being able to... Figure 1 Used in the example architecture;
[0016] Figure 6The illustration shows an example of run-length encoding (RLE), such as that that can be used in... Figure 1 Used in the example architecture;
[0017] Figure 7 The diagram shows that Figure 5 and Figure 6 Examples of combinations of compression schemes, such as those that can be used... Figure 1 This occurs in the example architecture;
[0018] Figure 8 The diagram illustrates an example workflow for cardinality clustering, such as... Figure 1 Used in the example architecture;
[0019] Figure 9 The diagram illustrates an exemplary architecture for cardinality clustering, such as that that can be used in... Figure 1 Used in the example architecture;
[0020] Figures 10-13 The instructions are shown when using, such as Figure 1 The example architecture is a flowchart illustrating exemplary operations that can be performed; and
[0021] Figure 14 A block diagram of an example computing device suitable for implementing some of the various examples disclosed herein is shown.
[0022] In all the accompanying drawings, the corresponding reference numerals indicate the corresponding parts. Detailed Implementation
[0023] The example solutions are used to read compressed data directly from files in persistent storage into an in-memory repository without decompression. Some examples transcode the compression scheme used by the file to the compression scheme used by the in-memory repository, while in other examples, the in-memory repository is compatible with the compression scheme used by the file. In some examples using transcoding, cardinality clustering is used to accelerate transcoding. Cardinality clustering minimizes cache misses, thereby improving the efficiency of memory access. When responding to queries, these methods significantly improve cold start time.
[0024] The example solutions described in this paper reduce computational overhead by avoiding decompression and recompression, including avoiding decompression / compression operations involving bit packing and run-length encoding (RLE), and improving cache locality. The example solutions described in this paper further reduce computational overhead by reducing memory pressure, such as by eliminating the need for decompression buffers. As described below, cache performance is improved by keeping data in a compressed state to make it more likely to fit into the local cache, and by improving cardinality clustering of cache hit probability. These computational savings reduce the amount of computational hardware required and the power consumption needed to maintain or even improve data query performance.
[0025] Examples demonstrate these advantageous technical performance benefits by loading compressed data into an in-memory repository without decompressing the compressed data, and / or by transcoding a stored compressed dictionary into an in-memory compressed dictionary based on cardinality clustering.
[0026] Various examples will be described in detail with reference to the accompanying drawings. Preferably, the same reference numerals are used in all the drawings to refer to the same or similar parts. References to specific examples and implementations in this disclosure are provided for illustrative purposes only and are not intended to limit all examples unless indicated otherwise.
[0027] Figure 1 An example architecture 100 is illustrated that advantageously provides the ability to read compressed data directly from file 200 into an in-memory repository 300 without decompression. In architecture 100, user terminal 102 uses cloud resources 110 to execute query 104 over network 1430 (e.g., a computer network) and responsively receives query results 106. Cloud resources 110 include a data lake 112 for persistent storage and a query engine 142, which can be decoupled and use separately scalable computing resources located on nodes different from those used by the data lake 112.
[0028] In some examples, query engine 142 is an analytical query engine that uses a vectorized query execution kernel 144, which is capable of performing queries on compressed data, such as query 104 on compressed data 302 in the in-memory repository 300. In some examples, query engine 142 supports business intelligence (BI) analytics services. The in-memory repository 300 and compressed data 302 are shown in more detail in Figure 3, and regarding... Figure 14 Network 1430 is described in more detail.
[0029] Reader 120 retrieves compressed data 202 from file 200 in data lake 112 and loads the data as compressed data 302 into in-memory repository 300. In some examples, transcoder 140 transcodes compressed data 202 to compressed data 302 without decompression. In some examples, cardinality clustering level 136 performs cardinality clustering on the data stream 134 of the data passed from reader 120 to in-memory repository 300 to accelerate the transcoding process. Transcoding is a digital-to-digital conversion from one encoding scheme to another, such as converting the data compression scheme(s) used in compressed data 202 to the data compression scheme(s) used in compressed data 302. This is necessary when the vectorized query execution kernel 144, which is compatible with the data compression scheme(s) used in compressed data 302, is incompatible with the data compression scheme(s) used in compressed data 202.
[0030] exist Figure 2 The document shows file 200 and compressed data 202 in more detail. Figures 4 to 7 The document shows examples of compression that can be used in compressed data 202 and compressed data 302, and regarding... Figure 8 and Figure 9 The cardinality clustering level 136 is described in more detail.
[0031] Data Lake 112 is shown as having multiple files in addition to file 200, such as file 200a and file 200b. In some examples, Data Lake 112 provides persistent storage capabilities with atomicity, consistency, isolation, and durability (ACID) properties, as well as time propagation and other features. Files 200, 200a, and 200b can be column-formatted (column-oriented) data files, such as Parquet, and / or optimized row-column (ORC) files using a column-block-specific compressed dictionary. A column block is a block of data for a specific column. A column block consists of a set of rows selected from row groups and is contiguous in the data file. A row group is a horizontal division of data into rows and has column blocks for each column in the dataset. This is about Figure 2 It is shown in more detail.
[0032] Reader 120 has a block reader 122 that reads block columns from file 200, and a page reader 124 that reads multiple pages within a block column. Two options are illustrated for reader 120. One option uses a set of callbacks: an RLE callback 126 for reading RLE data and a literal data callback 128 for reading literal data that is not RLE compressed (although the literal data can be compressed by other compression schemes, such as bit packing). In some examples, RLE callback 126 is vectorized. Figure 5 and Figure 6Further details about RLE relative to literal data are shown in the figure.
[0033] Another option is to use a stateful enumerator 132 instead of the RLE callback 126 and the literal data callback 128. The stateful enumerator 132 does not break down the data, but it knows its read state and can react differently to the input based on its state, iterating over the entire column content from the block of columns being read. For example, based on whether the read event occurred from a literal or an RLE section within file 200, the value being read is either a literal value or an RLE value (which has a value and a count).
[0034] Figure 2 Further details for file 200 are illustrated. In file 200, column data is stored across row groups. Several types of encoding can be used. For RLE encoding, in some examples, each column block has its own dictionary. File 200 is shown with two row groups, row group 210 and row group 220, although a different number of row groups can be used in some examples.
[0035] Compressed data 202 uses two compression schemes, one for each row group, such as storage compression scheme 204 for row group 210 and storage compression scheme 206 for row group 220. Since the compression schemes vary across different columns, storage compression scheme 204 has two subsets: one subset for column 250, which uses storage compression dictionary 210a from row group 210, and another subset for column 260, which uses storage compression dictionary 210b from row group 210. Similarly, storage compression scheme 206 has two subsets: one subset for column 250, which uses storage compression dictionary 220a from row group 220, and another subset for column 260, which uses storage compression dictionary 220b from row group 220. File 200 therefore has multiple storage compression dictionaries, including storage compression dictionaries 210a, 220a, 210b, and 220b.
[0036] Columns 250 and 260 each span row groups 210 and 220. Although two columns are shown, different numbers of columns can be used in some examples. As shown, columns 250 and 260 each have M rows, with a first subset of rows in row group 210 and the remainder in row group 220. In this example, a larger number of row groups would distribute the rows differently. In the illustrated example, column block 212a in row group 210 has some rows from column 250, column block 222a in row group 220 has additional rows from column 250, column block 212b in row group 210 has some rows from column 260, and column block 222b in row group 220 has additional rows from column 260.
[0037] Column block 212a is shown as having three pages, pages 214a, 216a, and 218a, and using storage compression dictionary 210a. Some examples use a different number of pages in the column block. Column block 212b is shown as having three pages, pages 214b, 216b, and 218b, and using storage compression dictionary 210b. Column block 222a is shown as having three pages, pages 224a, 226a, and 228a, and using storage compression dictionary 220a. Column block 222b is shown as having three pages, pages 224b, 226b, and 228b, and using storage compression dictionary 220b.
[0038] Figure 3A The diagram illustrates a comparison of the data structures between file 200 and in-memory repository 300. For illustrative purposes, corresponding columns across file 200 and in-memory repository 300 are shown. However, file 200 breaks down the rows of columns 250 and 260 into blocks within row groups, allowing for different compression dictionaries between row groups, while in-memory repository 300 uses a single compression dictionary per column. Some examples may use a single compression dictionary for the in-memory repository 300 as a whole.
[0039] In the illustrated example, column blocks 212a and 222a are combined (after possible transcoding) into a single column dataset 312a spanning rows of column 250. Column dataset 312a uses an in-memory compression dictionary 310a, which is the result of combining in-memory compression dictionaries 210a and 220a (after possible transcoding). Similarly, column blocks 212b and 222b are combined (after possible transcoding) into a single column dataset 312b spanning rows of column 260. Column dataset 312b uses an in-memory compression dictionary 310b, which is the result of combining in-memory compression dictionaries 210b and 220b (after possible transcoding). In the example where different compression dictionaries are used for different columns, in-memory compression scheme 304 has two subsets: one subset for column 250, using in-memory compression dictionary 310a; and another subset for column 260, using in-memory compression dictionary 310b.
[0040] Figure 3B The diagram illustrates the target Figure 3AMore details of the comparison are provided. Storage compression dictionary 210a and storage compression dictionary 220a are combined to form in-memory compression dictionary 310a. When the layout of in-memory repository 300 is incompatible with either storage compression schemes 204 and 206 or file 200, transcoding from storage compression schemes 204 and 206 to in-memory compression scheme 304 is required. Storage compression dictionaries 210a and 220a are combined to form in-memory compression dictionary 310a using transcoding (e.g., via transcoder 140), and column blocks 212a and 222a are transcoded, except for concatenation, to create column dataset 312a.
[0041] Examples of incompatibility include different dictionary identifiers (see...) Figure 5 The different counts, numbers, and other compression differences used to trigger an RLE (Repeated Leakage) for repeated values or characters, such as differences in bit packing and value encoding compression parameters. As an example of an RLE difference, a compression scheme might trigger an RLE when a value is encountered 10 times (see [link to RLE example]). Figure 6 Another compression scheme can trigger RLE when a value is encountered 20 times.
[0042] Figure 3C The illustration shows a comparison of the data structures between file 200 and in-memory repository 300 when in-memory repository 300 is compatible with both storage compression schemes 204 and 206, and storage compression schemes 204 and 206 are therefore reused as in-memory compression schemes 304 and 306 without transcoding. In this scenario, in-memory compression scheme 304 matches storage compression scheme 204, where in-memory compression dictionary 310a matches storage compression dictionary 210a, and in-memory compression dictionary 310b matches storage compression dictionary 210b. Additionally, in-memory compression scheme 306 matches storage compression scheme 206, in-memory compression dictionary 320a matches storage compression dictionary 220a, and in-memory compression dictionary 320b matches storage compression dictionary 220b. As a result, with... Figure 3A The scenarios requiring transcoding are different, and for these scenarios, transcoding is optional.
[0043] In this example, column dataset 312a spans only the rows of column block 212a, while column dataset 322a spans the rows of column block 222a. Similarly, in this example, column dataset 312b spans only the rows of column block 212b, while column dataset 322b spans the rows of column block 222b.
[0044] Figure 4The illustration shows an example of bit packing compression in bit packing compression scheme 400. Due to the standard integer size, column 402 of the original data uses 32 bits per row. However, by using the minimum number of bits required to store the integer value, column 404 uses only 10 bits. Removing common powers of 10 reduces the number of bits per column in column 406 to 7. For example, 590 / 10 = 59. By making all rows within column 406 relative to a minimum value (e.g., by subtracting the minimum value), column 408 uses only 6 bits per column. For example, 59 - 11 = 48.
[0045] A concept similar to the variation between columns 404 and 406 is value encoding, which can be used for compression in some examples. In value encoding, a non-integer floating-point number is multiplied by a common power of 10 to produce an integer, which can typically be represented with fewer bits than using floating-point data types (e.g., rational numbers). Value encoding advantageously allows for the direct aggregation of data values without decoding.
[0046] Figure 5 The illustration shows an example of dictionary compression in dictionary compression scheme 500. Column 502 of the original data text includes the string "Seattle" twice, "Los Angeles" once, and "Redmond" once. The compressed dictionary 504 has a dictionary identifier 506 as an integer value and a set of dictionary entry strings 508. Dictionary entry 510a has a dictionary identifier of 1 and the string "Seattle". Dictionary entry 510b has a dictionary identifier of 2 and the string "Los Angeles". Dictionary entry 510c has a dictionary identifier of 3 and the string "Redmond". Note that in some examples, the protocol uses a 1-based index, which can be used for the dictionary identifier, instead of a 0-based index.
[0047] Column 512 of the literal data (e.g., non-RLE compressed) lists the dictionary identifiers based on how the corresponding original data text appears, such as the two occurrences of the dictionary identifier value 1 at the end. Column 512 has the same number of rows as column 502, but because dictionary compression uses integer values 1, 2, and 3 instead of longer text strings, column 512 uses fewer bits per row.
[0048] RLE can be used with or without a compressed dictionary. Figure 6 The illustration shows an example of RLE without dictionary encoding, as shown in RLE compression scheme 600. For simplicity, Figure 6 Expand the original data text horizontally. Data text / value 602 is 3 A's, 4 B's, 5 C's, 6 A's, 5 B's, 4 C's, and 6 D's: "AAABBBBCCCCCAAAAAABBBBBCCCCDDDDDD".
[0049] The RLE string 604 without using a compressed dictionary has a data value followed by the number of repetitions: "A3B4C5A6B5C4D6". This scheme uses a threshold of 2 repetitions before triggering an RLE. An example of an RLE string generated from a threshold of 4 repetitions before triggering an RLE is "AAAB4C5A6B5C4D6", where "AAA" is the literal data and the remaining "4C5A6B5C4D6" is the RLE compressed data. A simple example of transcoding is converting "A3B4C5A6B5C4D6" to "AAAB4C5A6B5C4D6" / converting "AAAB4C5A6B5C4D6" to "A3B4C5A6B5C4D6".
[0050] The various compression techniques described in this article are complementary and can be used together. For example, RLE can be used with or without dictionaries, and with or without bit packing.
[0051] Figure 7 The illustration shows an example of combining dictionary compression scheme 500 and RLE compression scheme 600 into a combined data compression scheme 700. Uncompressed data 702 is used to generate a compressed dictionary 704 with dictionary identifier 706, which has values 1 to 4 (using 1-based indexing) and strings 708 that uniquely match the values (590, 110, 680, and 320) of the uncompressed data 702. Data row set 710 pairs dictionary identifier 706 with RLE counts 712 (e.g., repetitions of the same value). It can be seen that data row set 710 has fewer rows than uncompressed data 702. The compressed dictionary 704 and data row set 710 together form compressed data 714.
[0052] Note that in Figure 7 In the example shown, uncompressed data 702 is used with Figure 4 The same value shown in column 402. To use bit-packing compression scheme 400 with combined data compression scheme 700, the unique value of uncompressed data 702 in string 708 is only from... Figure 4 The bit-packed values of column 408 are replaced. This replaces the value set {590, 110, 680, 320} with {48, 0, 57, 21}. Since the data row set 710 has already been compressed, adding bit packing to the combined data compression scheme 700 only compresses the compression dictionary 704.
[0053] When creating a globally compressed dictionary for an in-memory repository (such as in-memory repository 300), typical approaches may include: (1) each thread using fine-grained lock-free operations or lightweight latches for each individual operation to perform element-wise lookup / insert / update calls at the hash bucket granularity; and (2) each thread acquiring a lock on the hash table and then running the lookup / insert / update operation while holding the lock. This technique can introduce high-performance overhead associated with memory and cache miss latency. This can be caused by random access patterns in the hash buckets, and by the scalability and contention issues that must be addressed with the concurrent access and synchronization required to leverage the expanded hash table size.
[0054] However, pre-sorting and / or clustering can make update insertions (a combination of updates and insertions) to the global compressed dictionary more scalable and cache-friendly. Partially reordering row values in row groups to populate the hash table and global dictionary maximizes cache locality because values that are close to each other in the hash table and global dictionary are inserted in the same order. This improves memory access efficiency and minimizes cache misses, minimizes the amount of time spent on random memory accesses and cache misses, minimizes resource contention (e.g., reducing the amount of time each thread holds locks or latches in the hash table structure), and reduces the CPU cache footprint (e.g., reducing cache line pollution).
[0055] Figure 8 The illustration depicts an example workflow 800 of cardinality clustering, which has been used in some examples to improve the speed of at least dictionary transcoding. In this illustrated example, file 200 has three line groups: line group 210 with at least a stored compressed dictionary 210a; line group 220 with at least a stored compressed dictionary 220a; and line group 230 with at least a stored compressed dictionary 230a. A dictionary read process 802a reads the stored compressed dictionary 210a from line group 210; a dictionary read process 802b reads the stored compressed dictionary 220a from line group 220; and a dictionary read process 802c reads the stored compressed dictionary 230a from line group 230.
[0056] Hash process 804a hashes at least a portion of the dictionary entries (e.g., dictionary identifiers) in the stored compressed dictionary 210a. After hashing by hash bucket identifiers, the hash values are sorted to provide optimal cache locality during dictionary insertion. Hash process 804b similarly hashes the dictionary entries in the stored compressed dictionary 220a, and hash process 804c similarly hashes the dictionary entries in the stored compressed dictionary 230a.
[0057] Procedure 806 updates hash table 908 (see also) Figure 9The input to the updated hash table 908a is sorted and clustered by hash buckets to ensure cache friendliness and provide optimal cache locality during dictionary insertion. In some examples, this reduces query latency by approximately 50% (e.g., from 22 seconds to 12 seconds). For example, update process 806a updates the hash table with the result from hash process 804a; update process 806b updates the hash table with the result from hash process 804b; and update process 806c updates the hash table with the result from hash process 804c.
[0058] Reading process 808a reads the RLE and literal data value from line group 210; reading process 808b reads the RLE and literal data value from line group 220; and reading process 808c reads the RLE and literal data value from line group 230. Finally, transcoding process 810a transcodes at least the storage compression dictionary 210a and the data read in reading process 808a; transcoding process 810b transcodes at least the storage compression dictionary 220a and the data read in reading process 808b; and transcoding process 810c transcodes at least the storage compression dictionary 230a and the data read in reading process 808c.
[0059] Figure 9 An exemplary architecture 900 for cardinality clustering is illustrated. Architecture 900 is illustrated at the stage where a stored compressed dictionary 210b is used to update hash table 908, which has previously been populated with hashed entries from the stored compressed dictionary 210a. Hash function 902 takes dictionary entries from the stored compressed dictionary 210b and generates hashed entries 904. Some of the hashed entries 904 are inserted into hash table 908 as insertion entries 906a, and some of the hashed entries 904 are duplicate entries 906b, and therefore do not need to be inserted. Duplicate entries 906b are those already in hash table 908, such as those from the previously processed compressed dictionary.
[0060] Architecture 900 sorts the vector of values to be inserted into hash table 908 by their associated hash values. This method provides a more efficient dictionary insertion access pattern, reducing cache misses. In some examples, sorter 910 uses a radix sort function. Radix sort is a non-comparison sorting algorithm that avoids comparisons by creating and assigning elements to buckets based on their radix. For elements with more than one significant digit, this bucketing process is repeated for each digit, maintaining the order of previous steps, until all digits have been considered. In some examples, radix sorting can be more beneficial for columns with medium to high radix. Inserting entry 906a into hash table 908 results in an updated hash table 908a.
[0061] Figure 10 A flowchart 1000 illustrating exemplary operations for software code vulnerability reduction that can be performed by architecture 100 is shown. In some examples, the operations described for flowchart 1000 are performed by... Figure 14 The computing device 1400 is used to execute the operation. Flowchart 1000 begins with receiving a query 104 (e.g., an input query) in operation 1002 (e.g., from network 1430). In some examples, receiving query 104 triggers a cold start in operation 1004.
[0062] Operation 1006 reads compressed data 202 from file 200. In some examples, reading compressed data 202 from file 200 includes reading file 200 from data lake 112. In some examples, reading compressed data from file 200 includes reading compressed data using a stateful counter 132, which distinguishes between literal data and RLE compressed data.
[0063] Operation 1008 loads compressed data 202 as compressed data 302 into the in-memory repository without decompressing compressed data 202. This is performed using decision operations 1010 through 1016. Decision operation 1010 determines the need for transcoding. If in-memory compression scheme 304 matches storage compression scheme 204 and storage compression scheme 206, and in-memory compression dictionary 310a matches storage compression dictionary 210a and storage compression dictionary 220a, then transcoding is not required, and flowchart 1000 moves to operation 1020.
[0064] Otherwise, in some examples, operation 1012 performs cardinality clustering of multiple stored compressed dictionaries (e.g., stored compressed dictionaries 210a and 220a). In some examples, this is based on... Figure 11 The process is executed using flowchart 1100.
[0065] Operation 1014 transcodes compressed data 202 from stored compression schemes 204 and 206 to in-memory compression scheme 304, for example, when the in-memory compression dictionary 310a differs from the stored compression dictionary 210a. Operation 1014 includes operations 1016 and 1018. Operation 1016 transcodes stored compression dictionaries 210a and 220a to the in-memory compression dictionary 310a. In some examples, this includes changing dictionary identifiers from those used in stored compression dictionaries 210a and 220a to those used in the in-memory compression dictionary 310a (see [link to example]). Figure 5In some examples, the transcoding of the storage compressed dictionaries 210a and 220a utilizes cardinality clustering from flowchart 1100. Operation 1018 transcodes the data values according to the transcoded compressed dictionary.
[0066] Operation 1020 executes query 104 on compressed data 302 in the in-memory repository 300, and operation 1022 returns query result 106. In some examples, executing query 104 includes using a vectorized query execution kernel 144.
[0067] Figure 11 A flowchart 1100 illustrating exemplary operations that can be performed by architecture 900 is shown. In some examples, the operations described for flowchart 1100 are performed by... Figure 14 The computational device 1400 is used to execute the process. Flowchart 1100 begins in operation 1002 with reading compressed data 202 from file 200, such as storing compressed dictionary 210a and (possibly in a second pass) storing compressed dictionary 220a. Operation 1104 hashes the entries in stored compressed dictionary 210a, and operation 1106 sorts the hashed entries.
[0068] Operation 1108 determines whether any hash entries need to be inserted into hash table 908. If not, flowchart 1100 proceeds directly to operation 1112. Otherwise, operation 1110 expands hash table 908 to insert the entry 906a, and updates hash table 908 in operation 1110.
[0069] Operation 1114 determines whether another compressed dictionary is included. If so, flowchart 1100 returns to operation 1102, and operation 1104 hashes the entries in another stored compressed dictionary (e.g., stored compressed dictionary 220a). When all compressed dictionaries have been processed, operation 1116 transcodes multiple stored compressed dictionaries (e.g., at least stored compressed dictionaries 210a and 220a).
[0070] Figure 12 A flowchart 1200 illustrating exemplary operations that can be performed by architecture 100 is shown. In some examples, the operations described for flowchart 1200 are performed by... Figure 14 The computing device 1400 is used to execute the operation. Flowchart 1200 begins with operation 1202, which includes reading compressed data from a file, the compressed data in the file having a first storage compression scheme, the first storage compression scheme having a first storage compression dictionary.
[0071] Operation 1204 includes loading compressed data into an in-memory repository without decompressing the compressed data. The compressed data in the in-memory repository has an in-memory compression scheme, and the in-memory compression scheme has an in-memory compression dictionary. Operation 1206 includes performing a query on the compressed data in the in-memory repository. Operation 1208 includes returning the query results.
[0072] Figure 13 A flowchart 1300 illustrating exemplary operations that can be performed by architecture 100 is shown. In some examples, the operations described for flowchart 1300 are performed by... Figure 14 The computing device 1400 is used to execute the operation. Flowchart 1300 begins with operation 1302, which includes reading compressed data from a file, the compressed data in the file having a first storage compression scheme, the first storage compression scheme having a first storage compression dictionary and a second storage compression dictionary.
[0073] Operation 1304 includes hashing the entries in the first and second stored compressed dictionaries. Operation 1306 includes sorting the hashed entries. Operation 1308 includes updating the hash table with the sorted hashed entries. Operation 1310 includes transcoding the first and second stored compressed dictionaries into an in-memory compressed dictionary, based at least on the sorted and updated hash table.
[0074] In another example, it is performed one at a time against a stored compressed dictionary. Figure 13 The operations illustrated are as follows. For example, hashing, sorting, updating, and transcoding are performed on one stored compressed dictionary, and then repeated on another stored compressed dictionary.
[0075] Additional examples
[0076] An example solution for directly reading compressed data into an in-memory repository involves reading compressed data from a file, where the compressed data in the file has a stored compression scheme with a stored compression dictionary. The compressed data is then loaded into the in-memory repository without decompressing it. The compressed data in the in-memory repository has an in-memory compression scheme. A query is performed on the compressed data in the in-memory repository, and the query results are returned. In some examples, the in-memory compression scheme differs from the stored compression scheme, while in other examples, the in-memory compression scheme and the stored compression scheme are the same. In some examples, the in-memory compression scheme uses a stored compression dictionary, and in others, the in-memory compression scheme uses an in-memory compression dictionary.
[0077] The example system includes: a processor; and a computer-readable medium storing instructions operable, when executed by the processor, to: read compressed data from a file having a first stored compression scheme with a first stored compression dictionary; load the compressed data into an in-memory repository without decompressing the compressed data, the compressed data in the in-memory repository having an in-memory compression scheme with an in-memory compression dictionary; perform a query on the compressed data in the in-memory repository; and return the query result.
[0078] An example computer-implemented method includes: receiving a query; reading compressed data from a file, the compressed data in the file having a first storage compression scheme, the first storage compression scheme having a first storage compression dictionary; loading the compressed data into a memory repository without decompressing the compressed data, the compressed data in the memory repository having a memory compression scheme, the memory compression scheme having a memory compression dictionary; performing a query on the compressed data in the memory repository; and returning the query result.
[0079] One or more example computer storage devices store computer-executable instructions that, when executed by a computer, cause the computer to perform operations including: reading compressed data from a file having a first storage compression scheme with a first storage compression dictionary; loading the compressed data into a memory repository without decompressing the compressed data, the compressed data in the memory repository having an in-memory compression scheme with an in-memory compression dictionary; receiving a query from a computer network; performing the query on the compressed data in the memory repository; and returning the query result.
[0080] Another example system includes: a processor; and a computer-readable medium storing instructions operable, when executed by the processor, for: reading compressed data from a file having a first stored compression scheme, the first stored compression scheme having a first stored compression dictionary and a second stored compression dictionary; hashing entries in the first and second stored compression dictionaries; updating the hash table with the hashed entries; sorting the updated hash table; and transcoding the first and second stored compression dictionaries into an in-memory compressed dictionary, at least based on the sorted updated hash table.
[0081] Another example computer-implemented method includes: reading compressed data from a file, the compressed data in the file having a first storage compression scheme, the first storage compression scheme having a first storage compression dictionary and a second storage compression dictionary; hashing entries in the first storage compression dictionary and the second storage compression dictionary; updating the hash table using the hashed entries; sorting the updated hash table; and transcoding the first and second storage compression dictionaries into an in-memory compression dictionary based at least on the sorted updated hash table.
[0082] One or more example computer storage devices have stored computer-executable instructions that, when executed by a computer, cause the computer to perform operations including: reading compressed data from a file having a first storage compression scheme with a first storage compression dictionary and a second storage compression dictionary; hashing entries in the first and second storage compression dictionaries; updating the hash table using the hashed entries; sorting the updated hash table; and transcoding the first and second storage compression dictionaries into an in-memory compression dictionary, at least based on the sorted updated hash table.
[0083] Alternatively, or in addition to the other examples described herein, examples include any combination of the following:
[0084] - Transcode compressed data from a storage compression scheme to an in-memory compression scheme;
[0085] -In-memory compression dictionaries differ from storage compression schemes;
[0086] - The first storage compression dictionary is specific to the first column and the first block of compressed data read from the file;
[0087] - The compressed data in the file also has a second storage compression dictionary;
[0088] - The second storage compression dictionary is specific to the first column and the second block of compressed data read from the file;
[0089] - The in-memory compression dictionary is the first column specific to the compressed data in the in-memory repository;
[0090] - The in-memory compression scheme matches the first storage compression scheme, and the in-memory compression dictionary matches the first storage compression dictionary;
[0091] - Reading compressed data from a file involves using a stateful enumerator that distinguishes between literal data and RLE compressed data to read the compressed data;
[0092] - Perform cardinality clustering of multiple storage compression dictionaries, including a first storage compression dictionary;
[0093] - Based at least on cardinality clustering, transcode multiple stored compressed dictionaries into an in-memory compressed dictionary;
[0094] - Hash the entries in multiple stored compressed dictionaries;
[0095] - Update the hash table using the hashed entries;
[0096] - Sort the updated hash table;
[0097] - Transcode multiple stored compressed dictionaries into in-memory compressed dictionaries, based at least on a sorted and updated hash table;
[0098] - Updating the hash table with hashed entries includes expanding the hash table for inserted entries;
[0099] - Sorting the updated hash table includes performing radix sort;
[0100] -Receive input queries;
[0101] - Receive input queries from a computer network;
[0102] - Receive input queries that trigger a cold start;
[0103] - The file includes a column-formatted file;
[0104] - The file uses the Parquet file format;
[0105] - Reading compressed data from a file includes reading files from a data lake;
[0106] - Reading compressed data from a file includes reading line groups;
[0107] - Reading compressed data from a file includes reading compressed data block by block;
[0108] - Reading compressed data from a file includes executing callbacks;
[0109] - Compressed data includes bit-packed data;
[0110] - Compressed data includes run-length encoded (RLE) compressed data;
[0111] - Transcoding compressed data includes transcoding a dictionary stored as a compressed dictionary into an in-memory compressed dictionary;
[0112] - Dictionary transcoding includes converting dictionary identifiers from those used in a first stored compressed dictionary to those used in an in-memory compressed dictionary.
[0113] - Dictionary transcoding also includes converting dictionary identifiers from those used in a second stored compressed dictionary to those used in an in-memory compressed dictionary.
[0114] - Load compressed data into an in-memory repository without packing the compressed data into bits;
[0115] - In response to receiving an input query, perform a query on the compressed data in the in-memory repository;
[0116] - Executing queries includes using a vectorized query execution kernel;
[0117] - The vectorized query execution kernel is compatible with in-memory compression schemes and uses an in-memory compression dictionary; and
[0118] - The vectorized query execution kernel is compatible with the first storage compression scheme and uses the first storage compression dictionary.
[0119] While aspects of this disclosure have been described in various examples of operations with their associations, those skilled in the art will understand that combinations of operations from any number of different examples are also within the scope of aspects of this disclosure.
[0120] Example operating environment
[0121] Figure 14 This is a block diagram of an example computing device 1400 (e.g., a computer storage device) used to implement the aspects disclosed herein, and is generally designated as computing device 1400. In some examples, one or more computing devices 1400 are provided to a local computing solution. In some examples, one or more computing devices 1400 are provided as a cloud computing solution. In some examples, a combination of local and cloud computing solutions is used. Computing device 1400 is merely one example of a suitable computing environment and is not intended to impose any limitation on the scope or functionality of the examples disclosed herein, whether used alone or as part of a larger set.
[0122] The computing device 1400 should not be interpreted as having any dependencies or requirements associated with any one or combination of the illustrated components / modules. The examples disclosed herein can be described in the general context of computer code or machine-usable instructions, including computer-executable instructions such as program components, executed by a computer or other machine such as a personal data assistant or other handheld device. Generally, a program component, including routines, programs, objects, components, data structures, etc., refers to code that performs a specific task or implements a specific abstract data type. The disclosed examples can be practiced in a variety of system configurations, including personal computers, laptops, smartphones, mobile tablets, handheld devices, consumer electronics, dedicated computing devices, etc. The disclosed examples can also be practiced in distributed computing environments when the task is performed by a remote processing device linked via a communication network.
[0123] Computing device 1400 includes a bus 1410 that directly or indirectly couples to the following devices: computer storage memory 1412, one or more processors 1414, one or more presentation components 1416, input / output (I / O) ports 1418, I / O components 1420, power supply 1422, and network components 1424. Although computing device 1400 is depicted as appearing as a single device, multiple computing devices 1400 can work together and share the depicted device resources. For example, memory 1412 can be distributed across multiple devices, and different devices can accommodate processor(s) 1414.
[0124] Bus 1410 represents a bus that can be one or more buses (e.g., an address bus, a data bus, or a combination thereof). Although lines are used for clarity. Figure 14 The various boxes can be used, but alternative representations can be used to complete the description of each component. For example, rendering components such as display devices are I / O components in some examples, and processors in some examples have their own memory. No distinction is made between categories such as "workstation," "server," "laptop," and "handheld device," because all these categories are considered to be within... Figure 14 Within the scope of, and collectively referred to herein as, a “computing device”. Memory 1412 may take the form of a computer storage medium as referenced below and is operatively provided to computing device 1400 with storage for computer-readable instructions, data structures, program modules, and other data. In some examples, memory 1412 stores one or more items from an operating system, a general-purpose application platform, or other program modules and program data. Therefore, memory 1412 is capable of storing and accessing data 1412a and instructions 1412b, which are executable by processor 1414 and configured to perform the various operations disclosed herein.
[0125] In some examples, memory 1412 includes computer storage media. Memory 1412 may include any amount of memory associated with or accessible by computing device 1400. Memory 1412 may be internal to computing device 1400 (e.g., Figure 14 The memory 1412 may be located outside the computing device 1400 (not shown), or both (not shown). Additionally or alternatively, the memory 1412 may be distributed across multiple computing devices 1400, for example, in a virtualized environment where instruction processing is performed on multiple computing devices 1400. For the purposes of this disclosure, "computer storage medium," "computer storage memory," "memory," and "memory device" are synonymous terms used for the memory 1412, and none of these terms include a carrier or propagation signaling.
[0126] Processor 1414 may include any number of processing units that read data from various entities such as memory 1412 or I / O components 1420. Specifically, processor(s) 1414 are programmed to execute computer-executable instructions for implementing aspects of this disclosure. These instructions may be executed by a processor, multiple processors within computing device 1400, or a processor external to client computing device 1400. In some examples, processor(s) 1414 are programmed to execute instructions such as those illustrated in the flowcharts discussed below and depicted in the accompanying drawings. Furthermore, in some examples, processor(s) 1414 represent implementations of analog techniques for performing the operations described herein. For example, the operations may be performed by analog client computing device 1400 and / or digital client computing device 1400. Multiple presentation components 1416 present data indications to a user or other device. Exemplary presentation components include display devices, speakers, printing components, vibration components, etc. Those skilled in the art will understand and recognize that computer data can be presented in a variety of ways, such as visually in a graphical user interface (GUI), audibly through speakers, wirelessly between computing devices 1400, across wired connections, or otherwise. I / O port 1418 allows computing device 1400 to be logically coupled to other devices including I / O components 1420, some of which may be built-in. Example I / O components 1420 include, for example, but not limited to, microphones, joysticks, gamepads, satellite antennas, scanners, printers, wireless devices, etc.
[0127] Computing device 1400 can operate in a networked environment via a logical connection to one or more remote computers through network component 1424. In some examples, network component 1424 includes a network interface card and / or computer-executable instructions (e.g., a driver) for operating the network interface card. Communication between computing device 1400 and other devices can occur via any wired or wireless connection using any protocol or mechanism. In some examples, network component 1424 is operable to use transport protocols via public, private, or hybrid (public and private) short-range communication technologies (e.g., Near Field Communication (NFC), Bluetooth) wirelessly. TM Communication data between devices (such as brand communications, etc.) or combinations thereof. Network component 1424 communicates with remote resource 1428 (e.g., cloud resource) across network 1430 via wireless communication link 1426 and / or wired communication link 1426a. Various examples of communication links 1426 and 1426a include wireless connections, wired connections, and / or dedicated links, and in some examples, are at least partially routed via the Internet.
[0128] While described in conjunction with example computing device 1400, the examples of this disclosure can be implemented with many other general-purpose or special-purpose computing system environments, configurations, or devices. Examples of well-known computing systems, environments, and / or configurations that may be applicable to various aspects of this disclosure include, but are not limited to, smartphones, mobile tablets, mobile computing devices, personal computers, server computers, handheld or laptop devices, multiprocessor systems, game consoles, microprocessor-based systems, set-top boxes, programmable consumer electronics, mobile phones, mobile computing and / or communication devices of wearable or accessory form factors (e.g., watches, glasses, headphones, or handsets), network PCs, minicomputers, mainframes, distributed computing environments including any of the systems or devices described above, virtual reality (VR) devices, augmented reality (AR) devices, mixed reality devices, holographic devices, etc. Such systems or devices can accept input from users in any manner, including input devices such as keyboards or pointing devices, gesture input, proximity input (such as by hovering), and / or voice input.
[0129] Examples of this disclosure can be described in the general context of computer-executable instructions (such as program modules) that are executed by one or more computers or other devices as software, firmware, hardware, or a combination thereof. Computer-executable instructions can be organized into one or more computer-executable components or modules. Generally, program modules include, but are not limited to, routines, programs, objects, components, and data structures that perform a particular task or implement a particular abstract data type. Aspects of this disclosure can be implemented with any number and organization of such components or modules. For example, aspects of this disclosure are not limited to the specific computer-executable instructions or specific components or modules illustrated in the figures and described herein. Other examples of this disclosure may include different computer-executable instructions or components having more or fewer functions than those illustrated and described herein. In examples involving general-purpose computers, aspects of this disclosure, when configured to execute the instructions described herein, transform a general-purpose computer into a special-purpose computing device.
[0130] By way of example and not limitation, computer-readable media include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable memory implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules. Computer storage media are tangible and mutually exclusive with communication media. Computer storage media are implemented in hardware and do not include carrier waves and propagating signals. Computer storage media used for the purposes of this disclosure are not signals themselves. Exemplary computer storage media include hard disks, flash drives, solid-state storage, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, optical disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage devices, magnetic tape cassettes, magnetic tape, disk storage devices or other magnetic storage devices, or any other non-transfer medium that can be used to store information for access by a computing device. In contrast, communication media typically embody computer-readable instructions, data structures, program modules, etc., in the form of modulated data signals such as carrier waves or other transmission mechanisms, and include any information delivery medium.
[0131] The order of execution or performance of the operations illustrated and described herein in the examples of this disclosure is not essential and may be performed in different orders in various examples. For example, it is conceivable that a particular operation is executed or performed before, simultaneously with, or after another operation within the scope of aspects of this disclosure. When introducing elements of aspects of this disclosure or its examples, the articles “a,” “an,” “the,” and “described” are intended to mean the presence of one or more of the elements. The terms “comprising / including” and “having” are intended to be inclusive and mean that additional elements may be present in addition to the listed elements. The term “exemplary” is intended to mean “an example of…”. The phrase “one or more of the following: A, B, and C” means “at least one of A and / or at least one of B and / or at least one of C.”
[0132] Various aspects of this disclosure have been described in detail, and it will be apparent that modifications and variations are possible without departing from the scope of the aspects of this disclosure as defined in the appended claims. Since various changes can be made to the foregoing constructions, products, and methods without departing from the scope of the aspects of this disclosure, it is intended that all content contained in the foregoing description and shown in the accompanying drawings be interpreted in an illustrative rather than restrictive sense.
Claims
1. A system comprising: processor; as well as A computer-readable medium storing instructions that, when executed by the processor, operate to: Compressed data is read from a file, wherein the compressed data in the file has a first storage compression scheme, and the first storage compression scheme has a first storage compression dictionary; Without decompressing the compressed data, the compressed data is loaded into an in-memory repository, the compressed data in the in-memory repository having an in-memory compression scheme, the in-memory compression scheme having an in-memory compression dictionary; Perform a query on the compressed data in the in-memory repository; as well as Return the query results.
2. The system according to claim 1, wherein the instructions are further operable to: The compressed data is transcoded from the first storage compression scheme to the in-memory compression scheme, wherein the in-memory compression dictionary is different from the first storage compression scheme.
3. The system according to claim 1, The first storage compression dictionary is specific to the first column and the first block of compressed data read from the file; The compressed data in the file also includes a second storage compression dictionary; The second storage compression dictionary is specific to the first column and the second block of compressed data read from the file; and The in-memory compression dictionary is specific to the first column of compressed data in the in-memory repository.
4. The system according to claim 1, wherein the in-memory compression scheme matches the first storage compression scheme, and wherein the in-memory compression dictionary matches the first storage compression dictionary.
5. The system of claim 1, wherein reading the compressed data from the file comprises reading the compressed data using a stateful enumerator or one or more callbacks that distinguish between literal data and run-length encoded (RLE) compressed data.
6. The system of claim 1, wherein the instructions are further operable to: Perform cardinality clustering of multiple storage compression dictionaries, including the first storage compression dictionary; and Based at least on the cardinality clustering, the plurality of stored compressed dictionaries are transcoded into the in-memory compressed dictionary.
7. The system of claim 6, wherein performing cardinality clustering of the plurality of stored compressed dictionaries comprises: Hash the entries in the plurality of stored compressed dictionaries; Update the hash table using the hashed entries; Sort the updated hash table; as well as The plurality of stored compressed dictionaries are transcoded into the in-memory compressed dictionary, based at least on the sorted and updated hash table.
8. A computer-implemented method, comprising: Receive queries; Compressed data is read from a file, wherein the compressed data in the file has a first storage compression scheme, and the first storage compression scheme has a first storage compression dictionary; Without decompressing the compressed data, the compressed data is loaded into an in-memory repository, the compressed data in the in-memory repository having an in-memory compression scheme, the in-memory compression scheme having an in-memory compression dictionary; The query is performed on the compressed data in the in-memory repository; and Return the query results.
9. The computer-implemented method according to claim 8, further comprising: The compressed data is transcoded from the first storage compression scheme to the in-memory compression scheme, wherein the in-memory compression dictionary is different from the first storage compression scheme.
10. The computer-implemented method according to claim 8, The first storage compression dictionary is specific to the first column and the first block of compressed data read from the file; The compressed data in the file also includes a second storage compression dictionary; The second storage compression dictionary is specific to the first column and the second block of compressed data read from the file; and The in-memory compression dictionary is specific to the first column of compressed data in the in-memory repository.
11. The computer-implemented method of claim 8, wherein the in-memory compression scheme matches the first storage compression scheme, and wherein the in-memory compression dictionary matches the first storage compression dictionary.
12. The computer-implemented method of claim 8, wherein reading the compressed data from the file comprises reading the compressed data using a stateful enumerator or one or more callbacks that distinguish between literal data and run-length encoded (RLE) compressed data.
13. The computer-implemented method according to claim 8, further comprising: Perform cardinality clustering of multiple storage compression dictionaries, the multiple storage compression dictionaries including the first storage compression dictionary; as well as Based at least on the cardinality clustering, the plurality of stored compressed dictionaries are transcoded into the in-memory compressed dictionary.
14. The computer-implemented method of claim 13, wherein performing cardinality clustering of the plurality of stored compressed dictionaries comprises: Hash the entries in the plurality of stored compressed dictionaries; Update the hash table using the hashed entries; Sort the updated hash table; as well as The plurality of stored compressed dictionaries are transcoded into the in-memory compressed dictionary, based at least on the sorted and updated hash table.
15. A computer storage device storing computer-executable instructions thereon, the computer-executable instructions causing the computer to perform operations when executed by a computer, the operations including: Compressed data is read from a file, wherein the compressed data in the file has a first storage compression scheme, and the first storage compression scheme has a first storage compression dictionary; Without decompressing the compressed data, the compressed data is loaded into an in-memory repository, the compressed data in the in-memory repository having an in-memory compression scheme, the in-memory compression scheme having an in-memory compression dictionary; Receive queries from a computer network; The query is performed on the compressed data in the in-memory repository; and Return the query results.
16. The computer storage device of claim 15, wherein the operation further comprises: The compressed data is transcoded from the first storage compression scheme to the in-memory compression scheme, wherein the in-memory compression dictionary is different from the first storage compression scheme.
17. The computer storage device according to claim 15, The first storage compression dictionary is specific to the first column and the first block of compressed data read from the file; The compressed data in the file also includes a second storage compression dictionary; The second storage compression dictionary is specific to the first column and the second block of compressed data read from the file; and The in-memory compression dictionary is specific to the first column of compressed data in the in-memory repository.
18. The computer storage device of claim 15, wherein the in-memory compression scheme matches the first storage compression scheme, and wherein the in-memory compression dictionary matches the first storage compression dictionary.
19. The computer storage device of claim 15, wherein reading the compressed data from the file comprises reading the compressed data using a stateful enumerator or one or more callbacks that distinguish between literal data and run-length encoded (RLE) compressed data.
20. The computer storage device of claim 15, wherein the operation further comprises: Hash the entries in a plurality of storage compression dictionaries, the plurality of storage compression dictionaries including the first storage compression dictionary; Update the hash table using the hashed entries; Sort the updated hash table; as well as The plurality of stored compressed dictionaries are transcoded into the in-memory compressed dictionary, based at least on the sorted and updated hash table.