Multi-domain configurable data compressor / decompressor
By using a shared registry to store compression operators in the data compressor and dynamically configure the compression pipeline, the problem of difficult to compress different data domains and structures in the prior art is solved, and efficient multi-domain data compression and decompression is achieved.
Patent Information
- Application Number
- CN202380070080.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-29
- Filing Date
- 2023-09-28
- Publication Date
- 2025-05-09
AI Technical Summary
The prior art is difficult to effectively compress data from different data domains or with different data structures, and the resources and costs of maintaining multiple domain-specific compression algorithms are high.
Using a configurable data compressor/decompressor, storing compression/decompression operators through a shared registry, dynamically configures compression pipelines to adapt to data from different data domains and structures.
It realizes efficient compression and decompression of data from multiple data domains and multiple data structures, reducing the cost and complexity of managing multiple compression algorithms.
Smart Images

Figure CN119968777A_ABST
Abstract
Description
Background Art
[0001] Data compression can be used to reduce the number of bits required to store or transmit data. Data compression algorithms are often developed with a specific data domain in mind. For example, a first type of compression algorithm may be designed to compress data from a first domain, such as general text data, while another type of compression algorithm may be designed to compress data from another domain, such as log data. Such domain-specific compression algorithms often exploit the redundancy of how information is structured in a given domain to achieve good compression performance for data in a given domain. However, a domain-specific compression algorithm designed for a first domain may not work well if used to compress data with different structures corresponding to different domains.
[0002] In addition, maintaining several different compression algorithms for a large number of data domains may be resource and cost intensive. For example, it may be necessary to manage updates of multiple compression algorithms. In addition, in some cases, determining which of several supported domain-specific compression algorithms to use for a given set of data to be compressed may not be straightforward. BRIEF DESCRIPTION OF THE DRAWINGS
[0003] Figure 1 is a block diagram illustrating a provider network including multiple cloud-based services according to some embodiments, wherein one of the services is a compressed data storage service including a configurable data compressor / decompressor configured to compress / decompress data from multiple data domains.
[0004] Figure 2 is a block diagram showing additional details about a compression recipe generation module including offline training, online training, hybrid training, and compression optimization according to some embodiments.
[0005] Figure 3 is a block diagram illustrating a configurable data compressor for compressing two different data sets using different recipes involving compression using different pipelines configured with compression operators stored in a common registry (eg, a pantry) according to some embodiments.
[0006] Figure 4 is a block diagram illustrating a configurable data decompressor for decompressing two different sets of data using different recipes involving decompression using different pipelines configured with decompression operators stored in a common registry (eg, a backup library), according to some embodiments.
[0007] Figure 5A diagram is shown representing an example compression / decompression pipeline that may be configured on a configurable compressor using compression operators stored in a registry, where the configurable compressor is configured to implement the compression / decompression pipeline based on a recipe for a given data type, according to some embodiments.
[0008] Figure 6 The relationship between the compressed file format, compressor / decompressor runtime, and recipes for the compressed file's data type is shown according to some embodiments.
[0009] Figure 7 is a flow chart illustrating corresponding processes for compressing and decompressing a data set using a data compressor / decompressor according to some embodiments.
[0010] Figure 8 is a flow chart illustrating an offline training process for learning recipes for different types of data sets according to some embodiments.
[0011] Fig. 9 is a flow chart illustrating an online training process for determining a recipe for a given data set, according to some embodiments.
[0012] Fig.10 is a flow chart illustrating a hybrid training process according to some embodiments, in which compressor operators considered during an online training process are learned from an offline training process to reduce the size of the compressor operators considered during the online training process.
[0013] Fig.11 is a flow chart illustrating the use of an optimization algorithm in determining a recipe for a given data type, according to some embodiments.
[0014] Fig.12 An example computer system is presented that can be used to implement an ontology-based annotation service and / or other systems or services as described herein.
[0015] Although embodiments are described herein by way of example for several embodiments and illustrative figures, those skilled in the art will recognize that embodiments are not limited to the described embodiments or drawings. It should be understood that the drawings and detailed descriptions thereof are not intended to limit the embodiments to the specific forms disclosed, but on the contrary, are intended to cover all modifications, equivalents and alternatives falling within the spirit and scope defined by the appended claims. The titles used herein are for organizational purposes only and are not intended to be used to limit the scope of this specification or claims. As used throughout this application, the word "may" is used in a permissible sense (i.e., meaning possible), rather than a mandatory sense (i.e., meaning must). Similarly, the words "include, including and includes" are meant to include but are not limited to. DETAILED DESCRIPTION
[0016] According to various embodiments, the systems and techniques described in the present disclosure implement a configurable data compressor / decompressor that can be adjustably configured to compress and decompress data from different data domains using a shared registry (e.g., a backup library of data compression / decompression operators).
[0017] As discussed above, data compression algorithms are often specifically designed to compress data belonging to a specific data domain or having a specific structure. Such custom-designed compression algorithms may provide good performance for data belonging to the data domain for which the data compression algorithm is designed, but may provide poor performance when used to compress data from other data domains or having other data structures. To address these limitations, some systems may use different data compression algorithms for different data domains. However, maintaining a large number of data compression algorithms may be difficult to manage, for example, and may require considerable cost and labor.
[0018] In some embodiments, a data storage service may include a configurable data compressor / decompressor that is capable of efficiently compressing and decompressing data from multiple data domains and having multiple data structures. For example, in some embodiments, the configurable data compressor / decompressor includes a shared compression / decompression operator registry (e.g., a backup library) and stores different recipes for compressing and decompressing different types of data sets. These recipes reference the compression / decompression operators of the registry (e.g., the backup library) and indicate the usage locations of the corresponding compression / decompression operators of the compression / decompression operators in the compression / decompression pipeline. For example, in some embodiments, a recipe may include a directed acyclic graph, in which the nodes of the graph indicate which compression / decompression operator in the registry (e.g., the backup library) will be deployed at a given node, and the edges of the graph show the data flow between the nodes as part of the compression or decompression pipeline.
[0019] In some embodiments, a data scientist may specify a recipe for a given data type, such as a customer or other user of a data storage service. Additionally or alternatively, in some embodiments, a compressed data storage service may include a compression recipe generation module configured to learn a corresponding recipe for compressing and decompressing data of a corresponding type (e.g., data belonging to different data domains). In some embodiments, offline training may be used to determine recipes for different types of data (e.g., different data domains). In some embodiments, offline training may involve using a graph including a plurality of compression operators at corresponding nodes of a graph. In some embodiments, the graph may be used to compress a set of training data, wherein the compression operator used at the corresponding node is iteratively changed to probe the performance of different combinations of compressor operators at the corresponding nodes of the graph. In such embodiments, the compression results of compressing the same training data using different combinations of compressor operators may be compared to select a recipe for a given type of data set (e.g., data domain). In addition, in some embodiments, the structure of a graph for performing offline learning may be determined by solving a graph optimization problem. In some embodiments, graph selection and compression operator selection may be performed simultaneously in a combined optimization problem. In some embodiments, a customer or other user may provide constraints and / or optimization goals for the graph optimization problem. For example, such constraints may include limits on memory to be used, limits on the amount of time to perform compression, limits on processing resources to be used, limits on power consumption, etc., and such optimization goals may include maximum compression ratio, minimum compression time, minimized resource usage, or other suitable optimization goals. In some embodiments, the first step in determining a recipe may be to determine a graph structure for the recipe, such as using a graph optimization problem. In addition, a second step may be to determine which compression operators to use at corresponding nodes of the selected graph, which may be determined using offline training as described above (or online or hybrid training as described below).
[0020] In some embodiments, online training may include partitioning the data set to be compressed into a plurality of blocks, and compressing corresponding blocks of the blocks using different combinations of compression operators at different nodes. In some embodiments, the graph structure used for online training may be selected from a set of stored graph structures, or may be selected based on a graph optimization problem using client / user constraints and objectives as described above. In online training, different combinations of compression operators may be used for consecutive blocks to be compressed, and compression results from previous blocks that have been compressed may be used to select a combination of compression operators to use for compressing subsequent blocks.
[0021] In some embodiments, a hybrid training process may be used, in which offline training is used to determine a subset of compression operators that are best suited for a given type of data (e.g., a data domain). In hybrid training, the available options for compression operators used on the corresponding nodes of a node may be truncated compared to a purely online training method, wherein the options are reduced to include those compression operators that are known to be best suited for the specific type of data being compressed. In some embodiments, the offline training method described above may be used to determine a subset of compression operators that are best suited for a given type of data (e.g., a data domain). In addition, in some embodiments, the hybrid training process may include using a recipe selected based on offline training as input and further modifying the recipe based on online training. For example, offline training may be used to determine a graph structure and an initial set of compressor operators for a given data type. In some embodiments of hybrid training, the graph and set of compressor operators determined by offline training may be used as a starting point for additional online training, in which the data to be compressed is separated into blocks, and the compressor operators are adjusted between compressions of corresponding blocks or otherwise adjusted to improve the starting configuration determined by offline training.
[0022] For example, consider a log file comprising a comma separated value (CSV) file having three fields: StartTime (string), EndTime (string), Metrics (consisting of multiple subfields). Figure 5 An example configurable data pipeline for such an example data domain input is shown in FIG, where offline, online, or hybrid training may have been used to generate Figure 5 The pipeline shown in FIG.
[0023] For example, in Figure 5 In the figure shown in , a "DataFrame" can be a columnized representation of a query log. The columns of the query log may include data representing "EndTime", "StartTime", and "Metrics". For example, these may be three fields encountered in a particular log. "Metrics" may further include subfields; therefore, the metrics field can be further parsed / split into corresponding subfields. The parsed fields (or subfields) can then be further transformed before being compressed with an entropy encoder to generate encoded data provided to the terminal node, where the compressed data provided to the terminal is written to the compressed output. For example, a conditional transformation can be applied to the end time, such as subtracting the start time from the end time so that the end represents the time after the start time, which may be a more compact data format than explicitly signaling the complete end time. Here, the start time is used as a side input of the conditional transformation, represented by a dotted arrow. During decompression, the start time needs to be decoded first before the end time can be decoded.
[0024] In some embodiments, a compressor runtime of a configurable compressor may read a recipe, load the transformations and dictionaries referenced by the recipe, and then actually compress the input objects, applying the transformations in the manner specified by the recipe (and possibly performing online learning to adapt to the data being compressed (e.g., online training or hybrid training). In some embodiments, the runtime writes an output file in a given format (e.g., .amzd). This format may reference a library of transformations required to decode it (e.g., a registry / backup library). However, the actual recipe used may not need to be noted in the output file. For example, the recipe may be inferred from characteristics of the output file, or the output file may embed a graph structure required for decompression.
[0025] Figure 1 is a block diagram illustrating a provider network including multiple cloud-based services according to some embodiments, wherein one of the services is a compressed data storage service including a configurable data compressor / decompressor configured to compress / decompress data from multiple data domains.
[0026] In some embodiments, the provider network 102 includes various cloud-based services, such as virtualized computing services 104, object-based storage services 106, block-based storage services 108, and other services 110. In addition, the provider network 102 includes a compressed data storage service 150. In some embodiments, the compressed data storage service 150 can compress "primary" data stored in other services of the provider network, such as data objects stored in the object-based storage service 106 or volume data stored in the block-based storage service 108. Additionally or alternatively, the compressed data storage service 150 can compress "auxiliary" data or "metadata" of various services, such as logs generated for the virtualized computing services 104, the object-based storage service 106, the block-based storage service 108, and the other services 110.
[0027] In some embodiments, the compressed data storage service 150 includes a compression recipe generation module 152 configured to generate a recipe for compressing a particular type of data set. For example, the compression recipe generation module 152 can use offline training, online training, hybrid training, graph optimization, etc. to select a recipe for compressing data of a particular data set type (e.g., a data domain). In some embodiments, a user (e.g., a data scientist) can provide instructions that specify a particular recipe for a particular data set type (e.g., a data domain).
[0028] In some embodiments, the compressed data storage service 150 also includes a compression recipe storage device 156, which can store information identifying corresponding recipes for data sets of different types (e.g., different data domains). In some embodiments, data to be compressed (or decompressed from) by the compressed data storage service 150 can be passed through a data input / output interface 158. In some embodiments, the client / user can indicate the data set type (e.g., data domain) through the data input / output interface 158. In addition, in some embodiments, the input / output interface 158 can infer the data set type based on the characteristics of the data to be compressed or based on information included in the file format of the compressed data to be decompressed. In some embodiments, the compressed data storage service 150 may include a storage device that stores compressed data within the compressed data storage service 150, or can provide compression and decompression services for data to be stored by other services (such as object-based storage services 106, etc.). In some embodiments, the object-based storage service 106 or other service 110 may provide virtualized storage resources to the compressed data storage service 150 to store data within the compressed data storage service 150 using virtualized storage resources that are physically implemented using storage devices of other services (such as the object-based storage service 106 or other service 110).
[0029] In some embodiments, the compression / decompression operator registry 154 (e.g., a reserve library of compression / decompression operators that can be used in recipes) stores various compression / decompression operators, such as data parsers, data splitters, data queue templates, various types of encoders, data transformations, conditional data transformations, context information shared resource templates, other compression operators, etc. In some embodiments, the compression / decompression operators stored in the compression / decompression operator registry 154 can be used at corresponding nodes in a plurality of nodes of a graph corresponding to a recipe for a given data set type (e.g., a data domain). For example, different recipes can mix and match the various compression / decompression operators stored in the compression / decompression operator registry 154 to generate customized compression / decompression pipelines for different data set types, wherein the selected compression / decompression operator selected for generating the customized compression / decompression pipeline is selected from a shared set of common compression / decompression operators stored in the compression / decompression operator registry 154.
[0030] In some embodiments, the configurable data compressor / decompressor 160 includes a software component 162 configured to determine a compression recipe for a given data set. For example, the data input / output interface 158 may provide an indication of the data set type, and the software component 162 may determine a recipe corresponding to the indicated data set type stored in the compression recipe storage device 156. The compressor / decompressor pipeline configuration module 164 may use the recipe determined by the software component 162 to configure a compression or decompression pipeline, the pipeline containing corresponding compression / decompression operators in the compression / decompression operator storage in the registry 154, where the operators are located at nodes of the graph according to the recipe, such as Figure 5 shown.
[0031] In addition, the configurable data compressor / decompressor 160 includes a runtime 166 that implements the configurable data compressor / decompressor 160 that has been dynamically configured using compression / decompression operators of a registry 154, wherein the registry is organized in a pipeline defined by a given recipe for compressing or decompressing data of a specific data set type (e.g., a data domain).
[0032] Figure 2 is a block diagram showing additional details about a compression recipe generation module including offline training, online training, hybrid training, and compression optimization according to some embodiments.
[0033] For example, in some embodiments, a compression recipe generation module of a compressed data storage service, such as a compression recipe generation module 152 of a compressed data storage service 150, may include an offline training component, an online training component, a hybrid training component, and an interface for receiving user-defined constraints and / or goals. For example, the compression recipe generation module 152 includes an offline training element 202, an online training element 204, a hybrid training element 206, and a storage element 208, which is configured to store user-provided constraints and / or optimization goals for offline, online, and / or hybrid training. For example, in some embodiments, the constraints and / or optimization goals may be provided as metadata included in a data set provided to the compressed data storage service 150 via the input / output interface 158. In some embodiments, the constraints and / or optimization goals may be provided via another interface, such as an administrator interface ( Figure 1 ), the administrator interface may allow a customer-designated administrator to define constraints and / or optimization goals for selecting a configurable compressor configuration for compressing data on behalf of the customer.
[0034] In some embodiments, offline training 202 may involve determining compressor elements stored in a registry 154 (e.g., a backup library) to be used at various nodes of the graph. For example, the graph may define a compressor pipeline to be used by the configurable compressor runtime 166 to implement a configurable compressor 160 that includes compressor elements indicated in a recipe at corresponding nodes of the graph. In some embodiments, graph structure selection and compressor element selection may be performed as part of offline training 202. In some embodiments, constraints and / or optimization goals stored in storage 208 may be used to perform offline training 202.
[0035] In some embodiments, online training 204 may similarly consider constraints and / or optimization goals stored in storage 208. In addition, in some embodiments, online training 204 may involve iterating multiple compressor element selections to compress individual blocks of multiple blocks of a data set to be compressed. For example, a first set of compressor elements may be used to compress a given block of a data set, and a different set of compressor elements may be used to compress another block of the data set. The compression performance of the first set of compressor elements and the different sets of compressor elements may be determined. This process may be repeated for any number of iterations. In some embodiments, the compression performance of a set of already considered compressor element combinations may be evaluated according to a threshold. In some embodiments, if one or more considered combinations satisfy one or more user-defined constraints and have a performance level within a threshold amount of one or more user-specified goals, the online training may be stopped, and a set of compressor elements that meet the criteria may be selected as a recipe for a given data type for which online training 204 is being performed. In some embodiments, additional combinations of compressor elements may continue to be evaluated as long as there are additional blocks of a data set to be compressed. In some embodiments, compressor element selection and / or parameters to be used for compressor elements may be adjusted for subsequent blocks to further improve performance.
[0036] In some embodiments, hybrid training 206 may involve taking an initial recipe determined using offline training 202 and further adjusting compressor element combinations and / or compressor element parameters when compressing corresponding blocks of a data set to further improve compression efficiency or other performance parameters. In addition, in some embodiments, based on information learned in offline training (such as offline training 202), some compressor elements and / or compressor element combinations may be excluded from additional consideration in hybrid training 206. For example, a recipe determined for a given data set type (e.g., data domain) learned using offline training 202 may exclude the use of certain compressor elements that are less suitable for the particular type of data set (e.g., data domain) being compressed. In some embodiments, excluding compressor elements that are less likely to be used can speed up the training performed in hybrid training 206.
[0037] Figure 3 is a block diagram illustrating a configurable data compressor for compressing two different data sets using different recipes involving compression using different pipelines configured with compression operators stored in a common registry (eg, a backup library) according to some embodiments.
[0038] In some embodiments, the same configurable data compressor 160 can implement various different compression pipelines 306 and 308 through the runtime 166 using various different recipes corresponding to different data set types (e.g., different data domains). For example, the configurable data compressor 160 can receive a first data set 302 to be compressed and a second data set 304 to be compressed. For example, the first data set 302 can include log data organized into three columns, such as a start time, an end time, and a metric (as previously discussed), and the second data set 304 can include a different configuration of the log data, such as a different number of columns or different types of data in the corresponding columns, or the second data set 304 can include completely different types of data, such as genomic data, as examples. In some embodiments, the data types of the corresponding data sets 302 and 304 can be transmitted to the configurable data compressor 160 (e.g., as metadata associated with the data sets 302 and 304), or in some embodiments, the configurable data compressor 160 can determine the corresponding data types of the data sets 302 and 304, for example, based on the structure and / or content of the data sets.
[0039] In some embodiments, component 162 of configurable data compressor 160 may determine the data type of the received data set, for example, based on associated metadata and / or by reasoning based on the structure and / or content of the received data set. Once the data set type is determined, a recipe corresponding to the determined data type may be retrieved from compression recipe storage 156, and compressor pipeline configuration module 164 may configure a set of executable program instructions including compressor elements referenced in the recipe from registry 154, which are arranged in a pipeline configuration specified in the recipe from compression recipe storage 156. The configured set of executable program instructions may then be executed in runtime 166 to implement a customized configuration of configurable data compressor 160, which has been configured to compress the corresponding data type of a given data set being compressed, such as data set 302 or 304.
[0040] In some embodiments, the output of the configurable data compressor 160 may include various terminal nodes from the data compression pipeline defined by the corresponding recipe (for example, Figure 5 In addition, output files 310 and 312 may include an indication of a graph representing a compression pipeline 306 or 308 for compressing data, and / or output files 310 and 312 may reference a respective recipe for compressing respective data sets 302 and 304 to generate output files 310 and 312, respectively.
[0041] For example, in some embodiments, a description of a pipeline graph 306 or 308 for decompressing a given one of the compressed output files 310 or 312 may be included in header data of the output file 310 or 312. In some embodiments, the header may reference a recipe for decompressing an output file such as the output file 310 or 312. In some embodiments, a description of a graph for implementing the pipeline 306 or 308 may be used in place of a reference to the recipe for the pipeline 306 or 308. In this way, the compressor and decompressor do not need to store the same recipe information, and this may allow for an independent decompressor that is independent of a compressor that compresses the decompressed data (e.g., the output file 310 or 312). In some embodiments, a third party may store recipe information that may be used by the configurable data compressor 160, wherein the third party provides the recipe information to the compressor side in order to generate the compressed file 310 or 312, and wherein the third party provides the recipe information to the decompressor side in order to decompress the compressed file 310 or 312.
[0042] Additionally, in some embodiments, a recipe may include multiple paths, where one path is selected over another path based on the data being compressed. In such cases, the compressed data file 310 or 312 may include one or more flags to indicate the path selected (or not selected) during compression.
[0043] Figure 4 is a block diagram illustrating a configurable data decompressor for decompressing two different sets of data using different recipes involving decompression using different pipelines configured with decompression operators stored in a common registry (eg, a backup library), according to some embodiments.
[0044] For example, the compressed version of Dataset 1 and the indicator of Recipe 1 (402) may be output 310 of Compression Pipeline 1 (306), such as Figure 3 Similarly, the compressed version of data set 2 and the indicator of recipe 2 (404) can be the output 312 of compression pipeline 2 (308), as shown in FIG. Figure 3 shown.
[0045] In some embodiments, a recipe indicator included in a compressed data file may be a description of a graph corresponding to a compression pipeline for compressing data included in the file. In some embodiments, a recipe indicator may be a description of a graph corresponding to a decompression pipeline for decompressing data included in the file. In some embodiments, a decompression pipeline for decompressing data included in a file may be different from a compression pipeline for compressing data included in the file. For example, a compression pipeline may have optional branches selected based on characteristics of the compressed data, while a decompression pipeline may omit optional branches that are not used in compression. In some embodiments, a recipe indicator included in a compressed data file may reference a recipe stored in a shared location without actually including a graph description in the compressed data file. For example, a compressed data file may reference a recipe stored in a shared location without actually including a graph description of the recipe in the compressed data file. In some embodiments, a compressor and a decompressor may maintain a set of similar compression operators that may be determined using recipe indicators. In some embodiments, a recipe indicator may be inferred based on the format of the compressed data in the compressed data file. For example, as Figure 5 As shown, each terminal node can write the compressed data as a separate compressed entity in the compressed data file. In some embodiments, the structure of the compression graph can be inferred based on the number, arrangement and encoding syntax of the compressed entities included in the compressed data file.
[0046] With Figure 3In a similar manner, the configurable data compressor / decompressor 160 may configure the configurable data compressor / decompressor 160 into a decompression pipeline using a recipe determined based on the recipe indicator in the compressed data files 402 and 404, as specified for the corresponding recipe corresponding to the compressed data files 404 and 402. In some embodiments, element 162 may be omitted because the recipe may be determined directly from the recipe indicator included in the compressed data file 402 or 404. In some embodiments, the compressor / decompressor pipeline configuration module 164 may compile the decompressor code based on the indicated recipe and using the compressor / decompressor operators from the registry 154 for execution in the runtime 166. For example, the decompression pipeline 406 may be implemented to decompress the compressed file 402 to generate data set 1 (410), and the decompression pipeline 408 may be implemented to decompress the compressed file 404 to generate data set 2 (412).
[0047] Figure 5 A diagram is shown representing an example compression / decompression pipeline that may be configured on a configurable compressor using compression operators stored in a registry, where the configurable compressor is configured to implement the compression / decompression pipeline based on a recipe for a given data type, according to some embodiments.
[0048] For example, the original log 502 may include a start time column, an end time column and a metric as a part of the original log file. The parser 504 may parse the original log file 502 into each column in the corresponding column. The output of the parser 504 may be stored in a data buffer, such as a buffer 510 (end time), a buffer 520 (start time) and a buffer 522 (metric). For example, conditional transformation 512 may be applied to the end time data to subtract the start time data. This may reduce the bit length of the end time data. In addition, the output of the conditional transformation may be stored in a buffer 514 (end time minus start time). In some embodiments, corresponding entropy encoders 516 and 526 may be used to encode the transformed end time data stored in the buffer 514 and the start time data stored in the buffer 520. In some embodiments, the context information sharing resource 530 may share context information between corresponding entropy encoders 516 and 526, or may reduce computational overhead by reusing context. The outputs of the respective entropy encoders 516 and 526 may be written to compressed data files via terminal nodes 518 and 528, where the terminal nodes contain only binary data that may be directly serialized into files. In some embodiments, the metrics stored in buffer 522 may be further split into different streams via splitter 524, which may be further processed using any compressor operator as shown in registry / backup library 500.
[0049] In some embodiments, the compressor / decompressor pipeline configuration module 164 can use the compression / decompression operators of the registry / repository 500 and configure the given recipe as specified in the given recipe according to the graph description of the given recipe. Figure 5 The pipeline shown (or any other pipeline). Note that the registry / backup library 500 includes example compression / decompression operators that can be used, but should not be interpreted as an exhaustive list. In some embodiments, various other compression / decompression operators can be used.
[0050] Figure 6 The relationship between the compressed file format, compressor / decompressor runtime, and recipes for the compressed file's data type is shown according to some embodiments.
[0051] In some embodiments, the compressor / decompressor runtime 602 (which may be similar to Figure 1 , 3 4) using a recipe 604 (which may be formatted as a JSON file or using another suitable file format) and a compressor operator included in a backup repository / registry 610. The recipe 604 may reference a compressor operator from the registry 610 to be used in the runtime 602. In addition, the output of the compressor may be a compressed file format 608 that may be read and written by the runtime 602 and reference the compressor operator of the backup repository / registry 610. For example, the compressed file format 608 may be decompressed in the runtime 602 using a decompressor operator from the backup repository 610, as indicated by a recipe indicator corresponding to the recipe 604. The decompressed file may then be read or written in the runtime 602 and later recompressed in the runtime 602 using a compressor operator from the backup repository / registry 610, as referenced by the recipe 604, to generate an updated compressed file 608. In some embodiments, the recipe 604 to be used may be generated by offline training, online training, hybrid training, or may be manually determined, such as by a data scientist.
[0052] Figure 7 is a flow chart illustrating corresponding processes for compressing and decompressing a data set using a data compressor / decompressor according to some embodiments.
[0053] At box 702, a configurable data compressor receives data to be compressed, such as log files, genomic data, text data, or other types of data. At box 704, a configurable compressor determines a compression recipe for compressing the received data. For example, the recipe can be indicated by a data scientist, or can be determined using offline, online, or hybrid learning, as described herein. At box 706, a configurable compressor dynamically configures itself (e.g., implementing a compression pipeline in the runtime of the compressor) based on the recipe determined at box 704 and using a compressor element of a backup library / registry accessible (or included therein) to the configurable compressor. At box 708, a configurable compressor compresses the received data in the compression runtime of the configurable compressor using the compression pipeline configured at box 706 based on the recipe determined at box 704. Then, at block 710, the configurable compressor provides a compressed version of the received data (e.g., a compressed data file) in a file format that indicates a description of a graph or indicates a recipe for configuring a decompression pipeline to decompress the decompressed version of the received data. In some embodiments, the received data may be uncompressed data, or may be data compressed using a different compression technique to be further compressed. In some embodiments, the received data may include a mix of compressed and uncompressed data to be further compressed.
[0054] To decompress the compressed version of the received data, at block 752, the configurable decompressor receives the compressed data to be decompressed and a graphical indication of a compression pipeline (or an indication of a recipe) for decompressing the compressed data. At block 754, the configurable decompressor dynamically configures itself (e.g., implements a decompression pipeline in the decompressor's runtime) based on a graph definition included in the compressed data file being decompressed or based on a recipe indicated by the compressed data file. In some embodiments, the graph definition and / or recipe may be inferred from the structure and / or syntax of the compressed data file. At block 756, the configurable decompressor decompresses the received compressed data at the decompression runtime of the configurable decompressor using the decompression pipeline configured at block 754 based on the recipe indicated in the compressed data file or otherwise provided with the compressed data file. Then, at block 758, the configurable decompressor provides the decompressed version of the received compressed data.
[0055] Figure 8 is a flow chart illustrating an offline training process for learning recipes for different types of data sets according to some embodiments.
[0056] At box 802, the offline training element receives training data of multiple data types. For example, each training data set may have a known data type, such as log data, genomic data, etc. In addition, as an example, the data type may be more fine-grained to distinguish between different types of logs. At box 804, the offline training element selects the first (or next) data set type to be used in training to determine the recipe to be used with the corresponding data set type. For example, to give a few examples, offline training can learn that some recipes are good at compressing log data, while other recipes are good at compressing genomic data. As another example, offline training can determine that some recipes are better at compressing certain types of log files, while other recipes are better at compressing other types of log files.
[0057] In some embodiments, constraints and optimization goals may be provided by a user, administrator, data scientist, etc., and may be used in offline training. For example, in some embodiments, compressor performance may be determined within a provided set of constraints, such as constraints on memory usage, processor usage, power usage, compression time, etc. Additionally, in some embodiments, compressor performance may be evaluated using optimization goals (e.g., achieving a high compression ratio, compressing in a minimum amount of time, conserving power, etc.). In some embodiments, optimization goals may be different from constraints, in that constraints may need to be satisfied, whereas optimization goals are used for evaluation, but no explicit constraint requirements are set.
[0058] Furthermore, in some embodiments, offline training as described herein may be performed on the configurable compressor as a whole (e.g., globally), or may be performed on corresponding subcomponents of the configurable compressor, such as branch-by-branch or node-by-node of a compression graph. For example, at block 806, a graph (or subcomponent of a graph) is generated (or selected) to be used in offline training, wherein the graph (or subcomponent of the graph) includes one or more nodes, and wherein at least one of the one or more nodes includes a set of compressor operators to be evaluated at the corresponding node.
[0059] At box 808, the offline training element uses a first combination of compressor operators from one or more sets of compressor operators to be evaluated at the corresponding node to compress the selected data set type as determined at box 804 (and / or box 806) (using a complete compressor or a portion of a compressor corresponding to the portion of the graph being evaluated). In addition, at box 810, the offline training element uses a second (or subsequent) combination of compressor operators from one or more sets of one or more compressor operators to be evaluated at the corresponding node to compress the same selected data set type as determined at box 804 (and / or box 806) (using the same complete compressor or the same portion of a compressor corresponding to the portion of the graph being evaluated). At box 812, the offline training element compares the performance of the corresponding compressor element combinations and determines whether a compressor element configuration can be assigned to the recipe or whether additional combinations need to be evaluated. In some embodiments, all compressor element combinations can be evaluated. In some embodiments, if the compressor performance meets one or more thresholds (which can be determined based on user-provided constraints and / or optimization goals), the offline training element can abandon evaluating additional combinations. In some embodiments, a predetermined number of combinations may be evaluated, and a given combination may be selected that has better compressor performance than other combinations. In some embodiments, different selection criteria may be used for different portions of the graph of the configurable compressor. For example, as an example, the selection criteria used at the branch level may be different from the selection criteria used at the global level. At block 814, the offline training element selects a compressor operator combination for the complete compressor pipeline graph (or portion thereof being evaluated) based on the compression results determined for the combinations considered in the loop including blocks 808, 810, and 812.
[0060] At block 816, the offline training element determines whether there are other types of data sets that need to be trained to determine the recipe, such as other types of log data, or genomic data, or various other data types in addition to the data type for which the recipe has been determined. If there are other types of data sets, the process returns to block 804, and if there are no other types of data sets, the training process ends at block 818. In some embodiments, the recipe and corresponding data set type associated with the determined recipe can be stored in the compression recipe storage device 156 and can be used to compress future data sets that are determined to have the same data set type as the data set type associated with the recipe.
[0061] Fig. 9 is a flow chart illustrating an online training process for determining a recipe for a given data set, according to some embodiments.
[0062] In some embodiments, online training may be performed in a similar manner to offline training. However, in online training, training may be performed "online" while compressing an actual data set submitted to a configurable compressor as the compression job to be performed, rather than using a training data set. In online training, the received data set may be divided into blocks, and the training results of the first block may be used to improve the configurable compressor configuration used to compress subsequent blocks.
[0063] For example, at block 902, the configurable compressor receives a data set to be compressed, and decomposes the received data set to be compressed into a plurality of data blocks at block 904. At block 906, the configurable compressor selects a graph (or a subcomponent of a graph) to be used in online training, wherein the graph (or a subcomponent of the graph) includes one or more nodes, and wherein at least one of the one or more nodes includes a set of compressor operators to be evaluated at the corresponding node.
[0064] At block 908, the online training element compresses the selected data block as determined at block 906 (using the full compressor or the portion of the compressor corresponding to the portion of the graph being evaluated) using a first combination of compressor operators from one or more sets of compressor operators to be evaluated at the corresponding node. Additionally, at block 910, the online training element compresses the next selected data block as determined at block 906 (using the same full compressor or the same portion of the compressor corresponding to the portion of the graph being evaluated) using a second (or subsequent) combination of compressor operators from one or more sets of compressor operators to be evaluated at the corresponding node, the next selected data block being similar in size to the first selected data block. At block 912, the online training element compares the performance of the corresponding compressor element combinations and determines whether the compressor element configurations can be updated in the online training of the current recipe or whether additional combinations need to be evaluated when compressing subsequent data blocks. In some embodiments, each data block can be compressed using compressor elements or parameters adjusted based on the compression performance of the previous one or more data blocks. In some embodiments, if the compressor performance meets one or more thresholds (which may be determined based on user-provided constraints and / or optimization goals), the online training element may forgo evaluating additional combinations. In some embodiments, online training may be performed continuously to allow the configurable compressor to adapt to changes in the structure of the data set over time. For example, the logarithmic condition may be dynamic, so that continuous adjustments may improve the performance of the compressor as the logarithmic characteristics change. At block 914, the online training element selects a compressor operator combination for the complete compressor pipeline graph (or the portion thereof being evaluated) to be used for the subsequent data block based on the compression results determined for the combinations considered in the loop including blocks 908, 910, and 912.
[0065] Fig.10 is a flow chart illustrating a hybrid training process according to some embodiments, in which compressor operators considered during an online training process are learned from an offline training process to reduce the size of the compressor operators considered during the online training process.
[0066] In some embodiments, the hybrid training process can be similar to the online training process. However, the hybrid training process can use an already trained recipe generated by offline training (or previous online or hybrid training) as the starting recipe for training. In addition, in some embodiments, in the hybrid training, the compressor operator options to be evaluated at the corresponding node can be truncated based on the data type. For example, at box 1002, the data set type can be determined, and at box 1004, at least some nodes based on the determined data set type can include only a subset of the available compressor operators. For example, based on the data set type, other compressor operators may be excluded from consideration.
[0067] Fig.11 is a flow chart illustrating the use of an optimization algorithm in determining a recipe for a given data type, according to some embodiments.
[0068] In some embodiments, constraints (as described herein) may be received at block 1102, and optimization goals (as described herein) may be received at block 1104. In some embodiments, a graph optimization problem may be solved at block 1106 to determine a graph for a recipe based on the constraints and optimization goals received at 1102 and 1104. Furthermore, in some embodiments, both the graph structure and the compressor operator selection may be varied iteratively, such as in Figure 8 In some embodiments, as described herein, constraints and optimization objectives can be used for offline training, online training, and / or hybrid training.
[0069] Fig.12 An example computer system, computer system 1200, is shown according to the discussed embodiments and examples, where computer system 1200 can be configured to implement a configurable compressor / decompressor, a compressed data storage service, other services, a resource host, a control plane, or other components of a service provider network service. In various embodiments, the computer system can be any of various types of devices, including, but not limited to, a personal computer system, a desktop computer, a laptop computer, a notebook or netbook computer, a mainframe computer system, a handheld computer, a workstation, a network computer, a mobile device, a consumer device, a video game console, a handheld video game device, an application server, a storage device, or generally any type of computing device or electronic device.
[0070] In addition, the systems and methods described herein can be implemented in various embodiments by any combination of hardware and software. For example, the method can be implemented by a computer system 1200 including one or more processors, and the one or more processors execute program instructions stored on a computer-readable storage medium coupled to the processor. The program instructions can be configured to implement the functions described herein (for example, the functions of various servers and other components of the data storage device described herein). The various methods shown in the accompanying drawings and described herein represent example embodiments of the method. The order of any method can be changed, and various elements can be added, reordered, combined, omitted or modified.
[0071] The computer system 1200 includes one or more processors 1210a-1210n (where any one of the processors may include multiple cores, which may be single-threaded or multi-threaded) coupled to a system memory 1220 via an input / output (I / O) interface 1230. The computer system 1200 further includes a network interface 1240 coupled to the I / O interface 1230. In various embodiments, the computer system 1200 may be a single processor system including one processor, or a multiprocessor system including several processors (e.g., two, four, eight, or another suitable number). The processor 1210 may be any suitable processor capable of executing instructions. For example, in various embodiments, the processor 1210 may be a general-purpose or embedded processor that implements any of a variety of instruction set architectures (ISAs), such as an x86, PowerPC, SPARC, or MIPS ISA, or any other suitable ISA. In a multiprocessor system, each of the processors 1210 may typically, but not necessarily, implement the same ISA. The computer system 1200 also includes one or more network communication devices (e.g., network interface 1240) for communicating with other systems and / or components via a communication network (e.g., the Internet, a LAN, etc.). For example, a client application executing on the system 1200 can use the network interface 1240 to communicate with a server application executing on a single server or server cluster that implements one or more of the components of the system described herein. In another example, an instance of a server application executing on the computer system 1200 can use the network interface 1240 to communicate with other instances of the server application (or another server application) that can be implemented on other computer systems. In addition, the computer system 1200 can be coupled to one or more input / output devices 1250, such as a cursor control device 1260, a keyboard 1270, a camera device 1290, and one or more displays 1280, via the I / O interface 1230.
[0072] In the illustrated embodiment, the computer system 1200 also includes one or more persistent storage devices and / or one or more I / O devices 1250. In various embodiments, the persistent storage device may correspond to a disk drive, a tape drive, a solid-state memory, other mass storage device, or any other persistent storage device. The computer system 1200 (or a distributed application or operating system running thereon) may store instructions and / or data in the persistent storage device as desired, and may retrieve the stored instructions and / or data as needed. For example, in some embodiments, the computer system 1200 may host a storage system server node, and the persistent storage device may include an SSD attached to the server node.
[0073] The computer system 1200 includes one or more system memories 1220 configured to store instructions and data accessible by the processor 1210. In various embodiments, the system memory 1220 may be implemented using any suitable memory technology, (e.g., one or more of the following: cache, static random access memory (SRAM), DRAM, RDRAM, EDO RAM, DDR 10 RAM, synchronous dynamic RAM (SDRAM), Rambus RAM, EEPROM, non-volatile / flash memory, or any other type of memory). The system memory 1220 may contain program instructions 1225 executable by the processor 1210 to implement the methods and techniques described herein. In various embodiments, the program instructions 1225 may be implemented in a platform native binary, such as a Java bytecode library, or a Java bytecode library. TM Any interpreted language such as byte code or C / C++, Java TM , or any other language or any combination thereof. For example, in the illustrated embodiment, program instructions 1225 include program instructions that can be executed to implement the compressed data storage service and / or configurable data compressor / decompression functions according to different embodiments. In some embodiments, program instructions 1225 can implement multiple separate clients, server nodes, and / or other components.
[0074] In some embodiments, program instructions 1225 may include instructions executable to implement an operating system (not shown), such as UNIX, LINUX, Solaris, TM 、MacOS TM , Windows TMAny of various operating systems such as 1200 and 1200. Any or all of the program instructions 1225 may be provided as a computer program product or software, which may include a non-transitory computer-readable storage medium having instructions stored thereon, which may be used to program a computer system (or other electronic device) to perform a process according to various embodiments. A non-transitory computer-readable storage medium may include any mechanism for storing information in a machine (e.g., computer) readable form (e.g., software, processing application). In general, a non-transitory computer-accessible medium may include a computer-readable storage medium or a memory medium, such as a magnetic medium or an optical medium, for example, a disk or DVD / CD-ROM coupled to the computer system 1200 via an I / O interface 1230. A non-transitory computer-readable storage medium may also include any volatile or non-volatile medium (such as RAM (e.g., SDRAM, DDR SDRAM, RDRAM, SRAM, etc.), ROM, etc.), which may be included in some embodiments of the computer system 1200 as a system memory 1220 or another type of memory. In other embodiments, program instructions may be transmitted using light, sound or other forms of propagated signals (e.g., carrier waves, infrared signals, digital signals, etc.) via communication media such as a network and / or wireless link, such as may be implemented via network interface 1240.
[0075] In some embodiments, system memory 1220 may include data storage 1235, which may be configured as described herein. In general, system memory 1220 (e.g., data storage 1235 within system memory 1220), persistent storage, and / or remote storage may store data blocks, copies of data blocks, metadata associated with data blocks and / or their states, configuration information, and / or any other information useful for implementing the methods and techniques described herein.
[0076] In one embodiment, I / O interface 1230 may be configured to coordinate I / O traffic between processor 1210, system memory 1220, and any peripheral devices in the system, including through network interface 1240 or other peripheral interfaces. In some embodiments, I / O interface 1230 may perform any necessary protocol, timing, or other data transformations to convert data signals from one component (e.g., system memory 1220) into a format suitable for use by another component (e.g., processor 1210). In some embodiments, for example, I / O interface 1230 may include support for devices attached through various types of peripheral buses, such as variants of the Peripheral Component Interconnect (PCI) bus standard or the Universal Serial Bus (USB) standard. In some embodiments, for example, the functionality of I / O interface 1230 may be split into two or more separate components, such as a north bridge and a south bridge. Moreover, in some embodiments, some or all of the functionality of I / O interface 1230 (such as the interface of system memory 1220) may be directly incorporated into processor 1210.
[0077] For example, the network interface 1240 can be configured to allow data to be exchanged between the computer system 1200 and other devices attached to the network, such as other computer systems (which can implement one or more virtual computing instances, temporary computing resources, storage system resource hosts, database engine head nodes or resource hosts, and / or clients of the systems described herein). In addition, the network interface 1240 can be configured to allow communication between the computer system 1200 and various I / O devices 1250 and / or remote storage devices. In some embodiments, the input / output device 1250 may include one or more display terminals, keyboards, keypads, touch pads, scanning devices, voice or optical recognition devices, or any other device suitable for inputting or retrieving data by one or more computer systems 1200. Multiple input / output devices 1250 may exist in the computer system 1200 or may be distributed on different nodes of a distributed system including the computer system 1200. In some embodiments, similar input / output devices may be separated from the computer system 1200 and may interact with one or more nodes of a distributed system including the computer system 1200 through a wired or wireless connection (such as through the network interface 1240). The network interface 1240 may typically support one or more wireless networking protocols (e.g., Wi-Fi / IEEE 802.11 or another wireless networking standard). However, in various embodiments, the network interface 1240 may support communications over any suitable wired or wireless general data network (e.g., other types of Ethernet). Additionally, the network interface 1240 may support communications over a telecommunications / telephone network (e.g., an analog voice network or a digital fiber optic communications network), over a storage area network (e.g., a Fiber Channel SAN), or over any other suitable type of network and / or protocol. In various embodiments, the computer system 1200 may include more, fewer, or different components than those shown (e.g., a display, a video card, a sound card, peripheral devices, other network interfaces, such as an ATM interface, an Ethernet interface, a frame relay interface, etc.).
[0078] It should be noted that any distributed system embodiment or any component thereof in the distributed system embodiment described herein can be implemented as one or more network-based services. For example, the computing cluster in the computing service can present computing services to the client and / or adopt other types of services of the distributed computing system described herein as network-based services. In some embodiments, network-based services can be implemented by software and / or hardware systems designed to support interoperable machines on the network to machine interactions. Network-based services can have interfaces described in machine-processable formats (e.g., Web Services Description Language (WSDL)). Other systems can interact with network-based services in a manner specified by the description of the interface of network-based services. For example, network-based services can define various operations that other systems can call, and can define specific application programming interfaces (APIs) that other systems can follow when requesting various operations.
[0079] In various embodiments, a network-based service may be requested or invoked using a message that includes parameters and / or data associated with the network-based service request. Such messages may be formatted according to a particular markup language such as Extensible Markup Language (XML), and / or may be encapsulated using a protocol such as Simple Object Access Protocol (SOAP). To execute a network-based service request, a network-based service client may assemble a message including the request using an Internet-based application layer transport protocol such as Hypertext Transfer Protocol (HTTP), and transmit the message to an addressable endpoint (e.g., a Uniform Resource Locator (URL)) corresponding to the network-based service.
[0080] In some embodiments, a network-based service may be implemented using a representational state transfer ("RESTful") technique rather than a message-based technique. For example, a network-based service implemented according to RESTful techniques may be invoked via parameters included in an HTTP method (such as PUT, GET, or DELETE) rather than encapsulated in a SOAP message.
[0081] Although the above embodiments have been described in considerable detail, numerous changes and modifications may be made, as will become apparent to those skilled in the art once the above disclosure is fully understood. The following claims are intended to be interpreted as covering all such modifications and changes, and the above description should therefore be regarded as illustrative rather than restrictive.
[0082] Embodiments of the present disclosure may be described in light of the following terms:
[0083] Clause 1. A system comprising:
[0084] a plurality of data storage devices; and
[0085] One or more computing devices, the one or more computing devices being configured to implement a data storage service, the data storage service comprising:
[0086] an interface configured to receive data to be stored in one or more data storage devices in the data storage device;
[0087] Data compression operator registry; and
[0088] A configurable data compressor, the configurable data compressor comprising:
[0089] A compressor configuration engine, the compressor configuration engine being configured to:
[0090] receiving or determining a first recipe for a first configuration of the configurable data compressor for compressing a first data set;
[0091] assembling a first compression pipeline comprising respective ones of the data compression operators in the registry arranged in a first configuration specified in the first recipe for the first data set; and
[0092] The compressor runtime is configured to:
[0093] compressing the first set of data using the configurable data compressor configured in the first configuration;
[0094] The compressor configuration engine is further configured to:
[0095] receiving or determining a second recipe for a second configuration of the configurable data compressor for compressing a second data set;
[0096] assembling a second compression pipeline comprising corresponding ones of the data compression operators in the registry arranged in a second configuration specified in the second recipe for the second data set; and
[0097] The compressor runtime is further configured as follows:
[0098] The second set of data is compressed using the configurable data compressor configured in the second configuration.
[0099] Clause 2. The system of clause 1, wherein the data storage service further comprises:
[0100] A configurable data decompressor, the configurable data decompressor comprising:
[0101] A decompressor configuration engine, the decompressor configuration engine being configured to:
[0102] receiving or determining the first formula;
[0103] assembling a first decompression pipeline, the first decompression pipeline comprising corresponding ones of the data compression operators in the registration table arranged in a first configuration specified in the first recipe; and
[0104] A decompressor runtime, the decompressor runtime being configured to:
[0105] decompressing a compressed version of the first set of data using the configurable data decompressor configured in the first configuration;
[0106] The decompressor configuration engine is further configured to:
[0107] receiving or determining the second formulation;
[0108] assembling a second decompression pipeline, the second decompression pipeline comprising corresponding ones of the data compression operators in the registration table arranged in a second configuration specified in the second recipe; and
[0109] The decompressor runtime is further configured to:
[0110] A compressed version of the second set of data is decompressed using the configurable data decompressor configured in the second configuration.
[0111] Clause 3. The system of clause 1 or clause 2, wherein the compression operators in the registry include two or more of:
[0112] Data parser;
[0113] Data queue;
[0114] A data encoder of a first type;
[0115] one or more additional types of data encoders;
[0116] compression transform; or
[0117] Compressed context information sharing connector.
[0118] Clause 4. A system according to any one of clauses 1 to 3, wherein the first data set or the second data set comprises:
[0119] Log Data; or
[0120] Genomic data.
[0121] Clause 5. The system of any one of clauses 1 to 4, wherein the data storage service further comprises:
[0122] A recipe generation module, wherein the recipe generation module is configured to:
[0123] Receiving training data of multiple types of data sets;
[0124] For the corresponding data set in the data set of the said type:
[0125] generating a graph for processing a data set of a given type, wherein the graph comprises a plurality of the data compression operators at corresponding ones of the nodes of the graph;
[0126] compressing the training data for a first time using a first set of corresponding data compression operators for the corresponding ones of the nodes of the graph;
[0127] compressing the training data one or more additional times using one or more additional sets of the corresponding data compression operators for the corresponding ones of the nodes of the graph; and
[0128] determining a recipe for the given type of data set based on performance of the corresponding data compression operator when compressing the training data the first time and the one or more additional times; and
[0129] Information is provided indicating that a corresponding one of the determined recipes is to be used for a corresponding one of the plurality of types of data sets.
[0130] Clause 6. The system of any one of clauses 1 to 4, wherein the data storage service further comprises:
[0131] A recipe generation module, wherein the recipe generation module is configured to:
[0132] sampling corresponding blocks of the first data set or the second data set;
[0133] generating a graph for processing the corresponding block of the first data set or the second data set, wherein the graph includes a plurality of the data compression operators at corresponding ones of the nodes of the graph;
[0134] compressing a first block using a first set of corresponding data compression operators for the corresponding ones of the nodes of the graph;
[0135] compressing one or more additional ones of the blocks using one or more additional sets of the corresponding data compression operators for the corresponding ones of the nodes of the graph; and
[0136] For a subsequent one of the blocks, the set of corresponding data compression operators is selected based on a performance metric of the data compression operators of a previous one of the blocks for use in compressing the subsequent one of the blocks.
[0137] Clause 7. The system according to clause 6, wherein the recipe generation module is further configured to:
[0138] Determine a data type of the first data set or the second data set, and reduce the plurality of data compression operators included at the corresponding ones of the nodes in the graph to include a reduced set of compression operators customized for the determined data type of the first data set or the second data set.
[0139] Clause 8. One or more non-transitory computer-readable storage media storing program instructions that, when executed on or across one or more processors, cause the one or more processors to:
[0140] receiving a first data set;
[0141] receiving or determining a first recipe for a first configuration of a configurable data compressor for compressing the first data set;
[0142] assembling a first compression pipeline, the first compression pipeline comprising respective ones of a plurality of data compression operators arranged in a first configuration specified in the first recipe;
[0143] compressing the first set of data using the configurable data compressor configured in the first configuration;
[0144] receiving a second data set;
[0145] receiving or determining a second recipe for a second configuration of the configurable data compressor for compressing the second data set;
[0146] assembling a second compression pipeline comprising respective ones of the plurality of data compression operators arranged in a second configuration specified in the second recipe; and
[0147] The second set of data is compressed using the configurable data compressor configured in the second configuration.
[0148] Clause 9. One or more non-transitory computer-readable storage media of Clause 8, wherein the program instructions, when executed on or across the one or more processors, further cause the one or more processors to:
[0149] providing a compressed version of the first data set, the compressed version comprising compressed data and an indicator of the compression operator used in the first compression pipeline;
[0150] A compressed version of the second set of data is provided, the compressed version comprising compressed data and an indicator of the compression operator used in the second compression pipeline.
[0151] Clause 10. One or more non-transitory computer-readable storage media according to clause 8 or clause 9, wherein the plurality of data compression operators are stored in a registry accessible by the one or more processors, wherein the plurality of data compression operators included in the registry include two or more of the following:
[0152] Data parser;
[0153] Data queue;
[0154] A data encoder of a first type;
[0155] one or more additional types of data encoders;
[0156] compression transform; or
[0157] Compressed context information sharing connector.
[0158] Clause 11. One or more non-transitory computer-readable storage media according to any one of clauses 8 to 10, wherein:
[0159] The first data set or the second data set includes log data.
[0160] Clause 12. One or more non-transitory computer-readable storage media according to any one of clauses 8 to 10, wherein the first data set or the second data set comprises two or more of the following:
[0161] Text data;
[0162] Image data;
[0163] Video data;
[0164] audio data; or
[0165] Genomic data.
[0166] Clause 13. One or more non-transitory computer-readable storage media according to any one of clauses 8 to 12, wherein the program instructions, when executed on or across the one or more processors, further cause the one or more processors to:
[0167] receiving one or more constraints to compress the first data set or the second data set; and / or
[0168] receiving one or more compression priorities to compress the first data set or the second data set; and
[0169] A graph optimization problem is solved for the first data set or the second data set to determine the first recipe or the second recipe, wherein the graph optimization problem is solved using the one or more constraints and / or the one or more compression priorities.
[0170] Clause 14. One or more non-transitory computer-readable storage media according to any one of clauses 8 to 13, wherein the program instructions, when executed on or across the one or more processors, further cause the one or more processors to:
[0171] Receiving training data of multiple types of data sets;
[0172] For the corresponding data set in the data set of the said type:
[0173] generating a graph for processing a data set of a given type, wherein the graph comprises a plurality of the data compression operators at corresponding ones of the nodes of the graph;
[0174] compressing the training data for a first time using a first set of corresponding data compression operators for the corresponding ones of the nodes of the graph;
[0175] compressing the training data one or more additional times using one or more additional sets of the corresponding data compression operators for the corresponding ones of the nodes of the graph; and
[0176] determining a recipe for the given type of data set based on performance of the corresponding data compression operator when compressing the training data the first time and the one or more additional times; and
[0177] Information is provided indicating that a corresponding one of the determined recipes is to be used for a corresponding one of the plurality of types of data sets.
[0178] Clause 15. One or more non-transitory computer-readable storage media according to any one of clauses 8 to 13, wherein the program instructions, when executed on or across the one or more processors, further cause the one or more processors to:
[0179] sampling corresponding blocks of the first data set or the second data set;
[0180] generating a graph for processing the corresponding block of the first data set or the second data set, wherein the graph includes a plurality of the data compression operators at corresponding ones of the nodes of the graph;
[0181] compressing a first block using a first set of corresponding data compression operators for the corresponding ones of the nodes of the graph;
[0182] compressing one or more additional ones of the blocks using one or more additional sets of the corresponding data compression operators for the corresponding ones of the nodes of the graph; and
[0183] For a subsequent one of the blocks, the set of corresponding data compression operators is selected based on a performance metric of the data compression operators of a previous one of the blocks for use in compressing the subsequent one of the blocks.
[0184] Clause 16. One or more non-transitory computer-readable storage media of Clause 15, wherein the program instructions, when executed on or across the one or more processors, further cause the one or more processors to:
[0185] Determine a data type of the first data set or the second data set, and reduce the plurality of data compression operators included at the corresponding ones of the nodes in the graph to include a reduced set of compression operators customized for the determined data type of the first data set or the second data set.
[0186] Clause 17. A method comprising:
[0187] receiving a first data set;
[0188] receiving or determining a first recipe for a first configuration of a configurable data compressor for compressing the first data set;
[0189] assembling a first compression pipeline, the first compression pipeline comprising respective ones of a plurality of data compression operators arranged in a first configuration specified in the first recipe;
[0190] compressing the first set of data using the configurable data compressor configured in the first configuration;
[0191] receiving a second data set;
[0192] receiving or determining a second recipe for a second configuration of the configurable data compressor for compressing the second data set;
[0193] assembling a second compression pipeline comprising respective ones of the plurality of data compression operators arranged in a second configuration specified in the second recipe; and
[0194] The second set of data is compressed using the configurable data compressor configured in the second configuration.
[0195] Clause 18. The method of clause 17, wherein the first compression pipeline or the second compression pipeline comprises a plurality of parallel branches, and wherein compressing the first data set or the second data set comprises compressing the parallel branches in parallel using a plurality of processing threads.
[0196] Clause 19. The method according to clause 17 or clause 18, further comprising:
[0197] Receiving training data of multiple types of data sets;
[0198] For the corresponding data set in the data set of the said type:
[0199] generating a graph for processing a data set of a given type, wherein the graph comprises a plurality of the data compression operators at corresponding ones of the nodes of the graph;
[0200] compressing the training data for a first time using a first set of corresponding data compression operators for the corresponding ones of the nodes of the graph;
[0201] compressing the training data one or more additional times using one or more additional sets of the corresponding data compression operators for the corresponding ones of the nodes of the graph; and
[0202] determining a recipe for the given type of data set based on performance of the corresponding data compression operator when compressing the training data the first time and the one or more additional times; and
[0203] Information is provided indicating that a corresponding one of the determined recipes is to be used for a corresponding one of the plurality of types of data sets.
[0204] Clause 20. The method according to clause 17 or clause 18, further comprising:
[0205] sampling corresponding blocks of the first data set or the second data set;
[0206] generating a graph for processing the corresponding block of the first data set or the second data set, wherein the graph includes a plurality of the data compression operators at corresponding ones of the nodes of the graph;
[0207] compressing a first block using a first set of corresponding data compression operators for the corresponding ones of the nodes of the graph;
[0208] compressing one or more additional ones of the blocks using one or more additional sets of the corresponding data compression operators for the corresponding ones of the nodes of the graph; and
[0209] For a subsequent one of the blocks, the set of corresponding data compression operators is selected based on a performance metric of the data compression operators of a previous one of the blocks for use in compressing the subsequent one of the blocks.
Claims
1. A system comprising: a plurality of data storage devices; and One or more computing devices, the one or more computing devices being configured to implement a data storage service, the data storage service comprising: an interface configured to receive data to be stored in one or more of the data storage devices in the data storage device; and Data compression operator registry; A configurable data compressor, the configurable data compressor comprising: A compressor configuration engine, the compressor configuration engine being configured to: receiving or determining a first recipe for a first configuration of the configurable data compressor for compressing a first data set; assembling a first compression pipeline, the first compression pipeline comprising respective ones of the data compression operators in the registration table arranged in a first configuration specified in the first recipe for the first data set; as well as The compressor runtime is configured to: compressing the first set of data using the configurable data compressor configured in the first configuration; The compressor configuration engine is further configured to: receiving or determining a second recipe for a second configuration of the configurable data compressor for compressing a second data set; assembling a second compression pipeline comprising corresponding ones of the data compression operators in the registry arranged in a second configuration specified in the second recipe for the second data set; and The compressor runtime is further configured as follows: The second set of data is compressed using the configurable data compressor configured in the second configuration.
2. The system according to claim 1, wherein the data storage service further comprises: A configurable data decompressor, the configurable data decompressor comprising: A decompressor configuration engine, wherein the decompressor configuration engine is configured to: receiving or determining the first formula; assembling a first decompression pipeline, the first decompression pipeline comprising corresponding ones of the data compression operators in the registration table arranged in a first configuration specified in the first recipe; as well as A decompressor runtime, the decompressor runtime being configured to: decompressing a compressed version of the first set of data using the configurable data decompressor configured in the first configuration; The decompressor configuration engine is further configured to: receiving or determining the second formulation; assembling a second decompression pipeline, the second decompression pipeline comprising corresponding ones of the data compression operators in the registration table arranged in a second configuration specified in the second recipe; and The decompressor runtime is further configured to: A compressed version of the second set of data is decompressed using the configurable data decompressor configured in the second configuration.
3. The system according to claim 1 or claim 2, wherein the compression operators in the registry include two or more of the following: Data parser; Data queue; A data encoder of a first type; one or more additional types of data encoders; compression transform; or Compressed context information sharing connector.
4. The system according to any one of claims 1 to 3, wherein the first data set or the second data set comprises: Log Data; or Genomic data.
5. The system according to any one of claims 1 to 4, wherein the data storage service further comprises: A recipe generation module, wherein the recipe generation module is configured to: Receiving training data of multiple types of data sets; For the corresponding data set in the data set of the said type: generating a graph for processing a data set of a given type, wherein the graph comprises a plurality of the data compression operators at corresponding ones of the nodes of the graph; compressing the training data for a first time using a first set of corresponding data compression operators for the corresponding ones of the nodes of the graph; compressing the training data one or more additional times using one or more additional sets of the corresponding data compression operators for the corresponding ones of the nodes of the graph; as well as determining a recipe for the given type of data set based on performance of the corresponding data compression operator when compressing the training data the first time and the one or more additional times; and Information is provided indicating that a corresponding one of the determined recipes is to be used for a corresponding one of the plurality of types of data sets.
6. The system according to any one of claims 1 to 4, wherein the data storage service further comprises: A recipe generation module, wherein the recipe generation module is configured to: sampling corresponding blocks of the first data set or the second data set; generating a graph for processing the corresponding block of the first data set or the second data set, wherein the graph includes a plurality of the data compression operators at corresponding ones of the nodes of the graph; compressing a first block using a first set of corresponding data compression operators for the corresponding ones of the nodes of the graph; compressing one or more additional ones of the blocks using one or more additional sets of the corresponding data compression operators for the corresponding ones of the nodes of the graph; and For a subsequent one of the blocks, the set of corresponding data compression operators is selected based on a performance metric of the data compression operators of a previous one of the blocks for use in compressing the subsequent one of the blocks.
7. The system according to claim 6, wherein the recipe generation module is further configured to: Determine a data type of the first data set or the second data set, and reduce the plurality of data compression operators included at the corresponding ones of the nodes in the graph to include a reduced set of compression operators customized for the determined data type of the first data set or the second data set.
8. One or more non-transitory computer-readable storage media storing program instructions that, when executed on or across one or more processors, cause the one or more processors to: receiving a first data set; receiving or determining a first recipe for a first configuration of a configurable data compressor for compressing the first data set; assembling a first compression pipeline, the first compression pipeline comprising respective ones of a plurality of data compression operators arranged in a first configuration specified in the first recipe; compressing the first set of data using the configurable data compressor configured in the first configuration; receiving a second data set; receiving or determining a second recipe for a second configuration of the configurable data compressor for compressing the second data set; assembling a second compression pipeline, the second compression pipeline comprising respective ones of the plurality of data compression operators arranged in a second configuration specified in the second recipe; and The second set of data is compressed using the configurable data compressor configured in the second configuration.
9. The one or more non-transitory computer-readable storage media of claim 8, wherein the program instructions, when executed on or across the one or more processors, further cause the one or more processors to: providing a compressed version of the first data set, the compressed version comprising compressed data and an indicator of the compression operator used in the first compression pipeline; A compressed version of the second set of data is provided, the compressed version comprising compressed data and an indicator of the compression operator used in the second compression pipeline.
10. One or more non-transitory computer-readable storage media according to claim 8 or claim 9, wherein the plurality of data compression operators are stored in a registry accessible by the one or more processors, wherein the plurality of data compression operators included in the registry include two or more of the following: Data parser; Data queue; A data encoder of a first type; one or more additional types of data encoders; compression transform; or Compressed context information sharing connector.
11. One or more non-transitory computer-readable storage media according to any one of claims 8 to 10, wherein: The first data set or the second data set includes log data.
12. One or more non-transitory computer-readable storage media according to any one of claims 8 to 10, wherein the first data set or the second data set comprises two or more of the following: Text data; Image data; Video data; audio data; or Genomic data.
13. One or more non-transitory computer-readable storage media according to any one of claims 8 to 12, wherein the program instructions, when executed on or across the one or more processors, further cause the one or more processors to: receiving one or more constraints to compress the first data set or the second data set; and / or receiving one or more compression priorities to compress the first data set or the second data set; and A graph optimization problem is solved for the first data set or the second data set to determine the first recipe or the second recipe, wherein the graph optimization problem is solved using the one or more constraints and / or the one or more compression priorities.
14. One or more non-transitory computer-readable storage media according to any one of claims 8 to 13, wherein the program instructions, when executed on or across the one or more processors, further cause the one or more processors to: Receiving training data of multiple types of data sets; For the corresponding data set in the data set of the said type: generating a graph for processing a data set of a given type, wherein the graph comprises a plurality of the data compression operators at corresponding ones of the nodes of the graph; compressing the training data for a first time using a first set of corresponding data compression operators for the corresponding ones of the nodes of the graph; compressing the training data one or more additional times using one or more additional sets of the corresponding data compression operators for the corresponding ones of the nodes of the graph; as well as determining a recipe for the given type of data set based on performance of the corresponding data compression operator when compressing the training data the first time and the one or more additional times; and Information is provided indicating that a corresponding one of the determined recipes is to be used for a corresponding one of the plurality of types of data sets.
15. One or more non-transitory computer-readable storage media according to any one of claims 8 to 13, wherein the program instructions, when executed on or across the one or more processors, further cause the one or more processors to: sampling corresponding blocks of the first data set or the second data set; generating a graph for processing the corresponding block of the first data set or the second data set, wherein the graph includes a plurality of the data compression operators at corresponding ones of the nodes of the graph; compressing a first block using a first set of corresponding data compression operators for the corresponding ones of the nodes of the graph; compressing one or more additional ones of the blocks using one or more additional sets of the corresponding data compression operators for the corresponding ones of the nodes of the graph; and For a subsequent one of the blocks, the set of corresponding data compression operators is selected based on a performance metric of the data compression operators of a previous one of the blocks for use in compressing the subsequent one of the blocks.