Ordered entropy blocks for higher space reduction

By sorting and reorganizing data blocks using entropy metrics, the problem of low block file compression efficiency in the prior art is solved, and more efficient storage and computing resource utilization is achieved.

CN119917472APending Publication Date: 2025-05-02COHESITY INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410716905.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-10-30
Filing Date
2024-06-04
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

The prior art leads to wasting space utilization, compression efficiency and computing resources when processing compression into block files.

Method used

By sorting and reorganizing data blocks using entropy metrics, the efficiency of the compression algorithm is improved, thereby generating more efficient block files.

Benefits of technology

It achieves higher compression rates and more efficient storage space utilization, reducing the consumption of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119917472A_ABST
    Figure CN119917472A_ABST
Patent Text Reader

Abstract

The invention relates to an ordered entropy block for achieving higher space reduction. Techniques are described for creating more efficient chunk files by using entropy metrics. In some examples, a processing device may determine an entropy value for each of a plurality of data blocks to obtain a corresponding plurality of entropy values. In some examples, the processing device may reorganize the plurality of data blocks based on the corresponding plurality of entropy values to obtain a reorganized plurality of data blocks. In some examples, a processing device may compress the reorganized plurality of data chunks to obtain a compressed chunk file. In some examples, a processing device may store the compressed chunk file in place of the plurality of data chunks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a data platform for a computing system. Background Art

[0002] Data platforms that support computing applications need to perform a variety of tasks, including regularly repeated end-user tasks, background tasks, and overhead tasks, all of which support the direct or indirect goals of the end-user and the overall efficiency of the data platform. When data is compressed into highly random block files, space utilization, compression efficiency, and computing resources (e.g., CPU cycles) may be wasted. Blocks are fragments of information used in many file formats. A block file is a set of data in which multiple blocks are embodied. Each of the multiple blocks includes stored data. A block may contain multiple files, one file, a portion of a file, or portions of multiple files, depending on how such files are written to the data storage area. For example, a large file may consume an entire block, while multiple small files may fit within a single block. In another example, a very large file may consume more space than the space allocated for a single block, and therefore, a single very large file may span multiple blocks. In any case, once the data is "chunked" into multiple blocks, a block file may be generated from multiple blocks, or a compressed block file may be generated from multiple blocks. Blocks are sometimes referred to as "data blocks."

[0003] In general, each block may include a header that specifies parameters such as block type, size, etc. The block header is followed by a variable data portion that may be decoded using the parameters in the header. Decoding the variable data portion allows recovery of the underlying information corresponding to the file within the block. Blocks compressed into block files are often used for archival data and / or static data (data that is not frequently modified), but this is not a technical requirement for creating blocks from such files and compressing the blocks into block files. In addition, block files are often compressed to improve storage efficiency, but compression is not a technical requirement for using blocks or creating block files from multiple blocks. Summary of the invention

[0004] Various aspects of the present disclosure describe techniques for creating more efficient block files by using an entropy metric. Using such an entropy metric applied to a data platform creates opportunities for many other optimizations in the creation and management of block files within a file system. For example, using an entropy metric can support improved malware detection, security enhancements, machine learning classification of data, encryption effectiveness, and the like.

[0005] In the context of a data platform, entropy is a measure of randomness or disorder. Generally speaking, the greater the randomness, the greater the disorder, resulting in less efficient compression of information stored on the file system. When applied to a data platform, the use of entropy can promote greater compressibility of stored data by using one or more techniques described in this disclosure.

[0006] In some examples, the processing device may determine an entropy value for one or more data blocks. In these and other examples, the processing device may sort the data blocks in ascending order using the entropy value of each corresponding block. The processing device may compress each of the one or more ordered data blocks into a block file. For example, positioning each of the one or more data blocks adjacent to each other (e.g., next to each other) based on the entropy value of each block may allow a compression algorithm to achieve a greater compression ratio, and thus the compressed block file may be more efficiently stored by the file system.

[0007] In some examples, data blocks with similar entropy values ​​and located together create greater opportunities for the compression algorithm to find patterns, thereby further reducing the storage space of the compressed block file. In some examples, data blocks with similar entropy values ​​and located together enable the compression algorithm to adjust smaller length codes, thereby producing smaller compressed block files that consume less storage space than compressing randomly ordered block files.

[0008] In some instances, the processing device may migrate or reorganize data blocks with similar entropy values ​​into a block file, thereby increasing or decreasing the order of entropy. In some examples, the steady-state data may be reorganized into a block file having data blocks arranged in ascending or descending order of entropy. The processing device may compress the data blocks into a block file using at least one of the *.gzip, *.zip, *.xz and / or *.bzip2 compression schemes. Different machine learning and / or deep learning models may use different types of entropy to determine the entropy value of a block or file. In some examples, bit entropy and / or byte entropy are calculated. The processing device may use the calculated bit entropy and / or byte entropy values ​​to evaluate whether the block or file is encrypted and / or compressed. The processing device may use the calculated bit entropy and / or byte entropy values ​​to evaluate potential compressibility. In some examples, the processing device may use the calculated bit entropy and / or byte entropy values ​​to evaluate the probability of malware existing in the block or file. In some examples, the processing device may use the calculated bit entropy and / or byte entropy values ​​to generate a heat map representing the entropy of the block or file.

[0009] In one example, various aspects of the technology are directed to a method. The exemplary method may include determining, by a processing device of a data platform, an entropy value for each of a plurality of data blocks to obtain a corresponding plurality of entropy values. The method may include: reorganizing, by the processing device and based on the corresponding plurality of entropy values, the plurality of data blocks to obtain a reorganized plurality of data blocks. Continuing with this example, the method may compress, by the processing device, the reorganized plurality of data blocks to obtain a compressed block file. The exemplary method may also store, by the processing device, the compressed block file in place of the plurality of data blocks.

[0010] In another example, various aspects of the technology relate to a data platform having a storage system and a processing device having access rights to the storage device. In such an example, the processing device is configured to perform various operations. For example, the processing device is configured to determine an entropy value for each of a plurality of data blocks to obtain a corresponding plurality of entropy values. In such an example, the processing device is further configured to reorganize the plurality of data blocks based on the corresponding plurality of entropy values ​​to obtain a reorganized plurality of data blocks. Continuing with this example, the processing device is further configured to compress the reorganized plurality of data blocks to obtain a compressed block file. In this example of the data platform, the processing device is further configured to store the compressed block file through the storage system to replace the plurality of data blocks within the storage system.

[0011] In another example, various aspects of the technology relate to a computer-readable storage medium having instructions that, when executed, configure one or more processors to perform various operations. In such an example, the instructions, when executed, may configure one or more processors to determine an entropy value for each of a plurality of data blocks to obtain a corresponding plurality of entropy values. In this example, the instructions, when executed, may configure one or more processors to reorganize the plurality of data blocks based on the corresponding plurality of entropy values ​​to obtain a reorganized plurality of data blocks. Continuing with this example, the instructions, when executed, may configure one or more processors to compress the reorganized plurality of data blocks to obtain a compressed block file. In this example of a computer-readable storage medium, the instructions, when executed, may configure one or more processors to store a compressed block file to replace the plurality of data blocks. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is a block diagram illustrating an example system for sorting block files according to entropy to achieve greater space reduction, in accordance with one or more techniques of this disclosure.

[0013] Figure 2 is a block diagram illustrating another example system that provides a more efficient data storage environment with data deduplication, encryption, and entropy thresholds, in accordance with one or more techniques of this disclosure.

[0014] Figure 3 is a block diagram illustrating an example system in accordance with techniques of this disclosure.

[0015] Figure 4 is a flow chart illustrating an example mode of operation of a computing device to create a more efficient block file by using an entropy metric, in accordance with the techniques of this disclosure.

[0016] Like reference numerals refer to like elements throughout the text and drawings. DETAILED DESCRIPTION

[0017] Figure 1is a block diagram illustrating an example system for sorting block files according to entropy to achieve greater space reduction in accordance with one or more techniques of this disclosure. Figure 1 In the example of , system 100 includes application system 102. Application system 102 represents a collection of hardware devices, software components, and / or data storage areas that can be used to implement one or more applications or services provided to one or more mobile devices 108 and one or more client devices 109 via network 113. Application system 102 may include one or more physical or virtual computing devices 172 that execute the computing workload of the application or service. Application system 102 includes a plurality of data blocks 174 having information to be stored. In Figure 1 In the example of , there are multiple data blocks 174, including each of data blocks 174A, 174B, 174C, and 174D (collectively referred to as "data blocks 174").

[0018] exist Figure 1 In the example of , the application system 102 includes application servers 170A to 170M (collectively referred to as "application servers 170") connected to a database server 172 that implements a database via a network. Database 172 may include one or more virtual machines, containers, Kubernetes pods each including one or more containers, bare metal processes, and / or other types of computing devices capable of performing work and storing information. Other examples of application system 102 may include one or more load balancers, web servers, network devices (such as switches or gateways), or other devices for implementing and delivering one or more applications or services to mobile devices 108 and client devices 109. Application system 102 may include one or more file servers. One or more file servers may implement the primary file system of application system 102. (In such instances, file system 153 may be a secondary file system that provides backup, archiving, and / or other services for the primary file system. References to file systems herein may include a primary file system or a secondary file system, for example, the primary file system of application system 102 or a file system 153 that operates as a primary file system or a secondary file system.)

[0019] The application system 102 may be located in a venue and / or in one or more data centers, each of which is part of a public cloud, a private cloud, or a hybrid cloud. The application or service may be a distributed application. The application or service may support enterprise software, financial software, office or other productivity software, data analysis software, end-user relationship management, web services, educational software, database software, multimedia software, information technology, healthcare software, or other types of applications or services. The application or service may be provided as software as a service (SaaS), platform as a service (PaaS), infrastructure as a service (IaaS), data storage as a service (dSaaS), or other types of services as a service (-aaS).

[0020] In some examples, application system 102 may represent an enterprise system, including one or more workstations in the form of desktop computers, laptops, mobile devices, enterprise servers, network devices, and other hardware to support enterprise applications. Enterprise applications may include enterprise software, financial software, office or other productivity software, data analysis software, end-user relationship management, web services, educational software, database software, multimedia software, information technology, healthcare software, or other types of applications. Enterprise applications may be delivered as a service from an external cloud service provider or other provider, executed locally on application system 102, or both.

[0021] exist Figure 1 In the example of , the system 100 includes a data platform 150 that provides a file system 153 and archiving functions to the application system 102 using the storage system 105 and the separate storage system 115. The data platform 150 implements a distributed file system 153 and a storage architecture to facilitate access to file system data by the application system 102 and to facilitate the transfer of data between the storage system 105 and the application system 102 via the network 111. In the case of a distributed file system, the data platform 150 enables the devices of the application system 102 to access the file system data via the network 111 using a communication protocol, as if such file system data were stored locally (e.g., stored on a hard disk of the device of the application system 102). Example communication protocols for accessing files and objects include Server Message Block (SMB), Network File System (NFS), or AMAZON Simple Storage Service (S3). The file system 153 can be a primary file system or a secondary file system for the application system 102.

[0022] The file system manager 152 represents a collection of hardware devices and software components that implement the file system 153 for the data platform 150. Examples of file system functions provided by the file system manager 152 include storage space management, including deduplication, file naming, directory management, metadata management, partitioning, and access control. The file system manager 152 executes communication protocols to facilitate application systems 102 to access files and objects stored in the storage system 105 via the network 111.

[0023] exist Figure 1 In the example of , system 100 includes a data platform 150 that provides a file system 153 and archiving functionality to application system 102 using storage system 105 and a separate storage system 115. Data platform 150 implements a distributed file system 153 and storage architecture to facilitate access to file system data by application system 102 and to facilitate data transfer between storage system 105 and application system 102 via network 111. As depicted herein, system storage 115 is represented as being co-located with data platform 150.

[0024] The data platform 150 includes a storage system 105 having one or more storage devices 180A-180N (collectively, "storage devices 180"). Storage devices 180 may represent one or more physical or virtual computers and / or storage devices that include or otherwise have access to storage media. Such storage media may include one or more of a flash drive, a solid state drive (SSD), a hard disk drive (HDD), an electrically programmable memory (EPROM), or an electrically erasable programmable (EEPROM) memory form factor and / or other types of storage media used to support the data platform 150. Different storage devices of storage devices 180 may have different types of storage media combinations.

[0025] In some examples, each of the storage devices 180 may include system memory. In some examples, each of the storage devices 180 may be a storage server, a network attached storage (NAS) device, or may represent disk storage of a computer device. The storage system 105 may be a redundant array of independent disks (RAID) system. In some examples, one or more of the storage devices 180 may be both a computing device and a storage device, which executes the software of the data platform 150, such as the file system manager 152 and the compression manager 154 in the example of the system 100, and stores the objects and metadata of the data platform 150 in the storage medium. In some examples, a separate computing device (not shown) may execute the software of the data platform 150, such as the file system manager 152 and the compression manager 154 in the example of the system 100. Each of the storage devices 180 may be considered and referred to as a "storage node" or simply a "node". The storage device 180 may represent a virtual machine running on a supported hypervisor, a cloud virtual machine, a physical rack server, or a computing model installed in a converged platform.

[0026] In some examples, data platform 150 may run locally on a physical system, virtually, or in a cloud. For example, data platform 150 may be deployed as a physical cluster, a virtual cluster, or a cloud-based cluster running in a private cloud, a hybrid private / public cloud, or a public cloud deployed by a cloud service provider. In some examples of system 100, multiple instances of data platform 150 may be deployed, and file system 153 may be replicated between the various instances. In some cases, data platform 150 is a computing cluster representing a single management domain. The number of storage devices 180 may be expanded to meet performance requirements.

[0027] In some examples, the data platform 150 may implement and provide multiple storage domains to one or more tenants, or to isolate workloads or storage requirements that require different data policies. A storage domain is a data policy domain that determines the policies for deduplication, compression, encryption, tiering, and other operations performed with objects stored on the storage domain. In this way, the data platform 150 may provide users with the flexibility to select global data policies or workload-specific data policies. The data platform 150 may support partitioning.

[0028] A view is a protocol export that resides within a storage domain. Although a view inherits data policies from its storage domain, additional data policies can be specified for a view. Views can be exported via SMB, NFS, S3, and / or another communication protocol. Policies that determine data processing and storage for the data platform 150 can be assigned at the view level. A protection policy can specify backup frequency and retention policies, which can include data lock periods. Archives 142 or snapshots created according to a protection policy inherit the data lock period and retention period specified by the protection policy.

[0029] Each of network 113 and network 111 may be the Internet or may include or represent any public or private communication network or other network. For example, network 113 may be a cellular, Near field communication (NFC), satellite, enterprise, service provider, and / or other types of networks that support the transfer of data between computing systems, servers, computing devices, and / or storage devices. One or more of such devices may transmit and receive data, commands, control signals, and / or other information across network 113 or network 111 using any suitable communication technology. Each of network 113 or network 111 may include one or more network hubs, network switches, network routers, satellite antennas, or any other network equipment.

[0030] Such network devices or components are operably coupled to each other, thereby providing for an exchange of information between computers, devices, or other components (eg, between one or more client devices or systems and one or more computer / server / storage devices or systems). Figure 1 Each of the devices or systems shown in the figure can be operatively coupled to the network 113 and / or the network 111 using one or more network links. The links coupling such devices or systems to the network 113 and / or the network 111 can be Ethernet, asynchronous transfer mode (ATM), or other types of network connections, and such connections can be wireless and / or wired connections. Figure 1 and Figure 2 One or more of the devices or systems shown in or on network 113 and / or network 111 may be remotely located relative to one or more other shown devices or systems.

[0031] The application system 102 may generate objects and other data using the file system 153 provided by the data platform 150, and the file system manager 152 may create, manage, and cause the objects and other data to be stored in the storage system 105. Therefore, the application system 102 may alternatively be referred to as a "source system", and the file system 153 of the application system 102 may alternatively be referred to as a "source file system". The application system 102 may communicate directly with the storage system 105 via the network 111 to transfer objects for some purposes, and may communicate with the file system manager 152 via the network 111 to indirectly obtain objects or metadata from the storage system 105 for some purposes.

[0032] In some examples, the file system manager 152 generates metadata and stores it to the storage system 105. The data set stored to the storage system 105 and used to implement the file system 153 is referred to herein as file system data. In some examples, the file system data may include data blocks 174 and block files 164 in various intermediate stages of compression, encryption, and storage. The compressed block files 176 may be stored within the storage system 105, 115 and / or stored via the archive 142. In some examples, the file system data may include block files 164 in the completed stages of compression, encryption, and storage. The file system data may include the metadata and objects mentioned above. The metadata may include file system objects, tables, trees, or other data structures; metadata generated to support deduplication; or metadata used to support snapshots. The stored objects may include files, virtual machines, databases, applications, pods, containers, any workloads, system images, directory information, or other types of objects used by the application system 102. Objects of different types and objects of the same type may be deduplicated with respect to each other. In some examples, the block files 164 replace the data blocks 174 in a lossless compressed format.

[0033] Aspects of the present disclosure describe techniques for creating more efficient block files 164 by using entropy metrics. Applying such entropy metrics to the data platform 150 creates opportunities for many other optimizations in the creation and management of block files 164 within the file system 153. For example, using entropy metrics can support improved malware detection, security enhancements, machine learning classification of data, encryption effectiveness, etc.

[0034] In the context of data platform 150, entropy is a measure of randomness or disorder. Generally speaking, greater randomness results in greater disorder, leading to less efficient compression of information stored on file system 153. When applied to data platform 150, the use of entropy can promote greater compressibility of stored data using one or more techniques described in this disclosure.

[0035] In some examples, the processing device 199 may determine an entropy value 186 for one or more data blocks 174. In these and other examples, the processing device 199 may sort the data blocks 174 in ascending order using the entropy value 186 for each corresponding block. The processing device 199 may compress each of the one or more ordered data blocks 174 into a block file 164. For example, positioning each of the one or more data blocks adjacent to each other (e.g., next to each other) based on the entropy value of each block may allow a compression algorithm to achieve a greater compression ratio, and thus the compressed block file 176 may be more efficiently stored by the file system 153.

[0036] In some examples, data blocks 174 having similar entropy values ​​186 and being located together create greater opportunities for the compression algorithm to find patterns, thereby further reducing the storage space of the compressed block file 176. In some examples, data blocks 174 having similar entropy values ​​186 and being located together enable the compression algorithm 107 to adjust smaller length codes, thereby generating a smaller compressed block file 176 that consumes less storage space than compressing a randomly ordered block file.

[0037] In some instances, the processing device 199 may migrate or reorganize data blocks 174 having similar entropy values ​​186 into a block file 164, thereby increasing or decreasing the order of entropy. In some examples, the steady-state data is reorganized into a block file 164 having data blocks arranged in ascending or descending order of entropy. The processing device 199 may compress the data blocks 174 into a block file using at least one of the *.gzip, *.zip, *.xz, and / or *.bzip2 compression algorithms 107. Different machine learning and / or deep learning models may utilize different types of entropy to determine the entropy value 186 of a data block 174 or file to be included in the block file 164. In some examples, a bit entropy value 186A and / or a byte entropy value 186B are calculated. The processing device 199 may use the calculated bit entropy 186A and / or byte entropy value 186B to evaluate whether the data block 174 or file is encrypted and / or compressed. Processing device 199 may use calculated bit entropy 186A values ​​and / or byte entropy values ​​186B to assess potential compressibility. In some examples, processing device 199 uses calculated bit entropy 186A values ​​and / or byte entropy values ​​186B to assess the probability of malware being present within data block 174 or file. In some examples, processing device 199 may use calculated bit entropy 186A values ​​and / or byte entropy values ​​186B to generate a heat map representing the entropy of data block 174 or file.

[0038] In one example, various aspects of the technology are directed to a method. The exemplary method may include determining, by a processing device 199 of a data platform 150, an entropy value 186 for each of a plurality of data blocks 174 to obtain a corresponding plurality of entropy values ​​186. The method may include: reorganizing, by the processing device 199 and based on the corresponding plurality of entropy values ​​186, the plurality of data blocks 174 to obtain a reorganized plurality of data blocks 174. Continuing with this example, the method may compress, by the processing device 199, the reorganized plurality of data blocks 174 to obtain a compressed block file 176. The exemplary method may also store, by the processing device 199, the compressed block file 176 in place of the plurality of data blocks 174.

[0039] In another example, various aspects of the technology relate to a data platform 150 having a processing device 199, a storage system 105, 115, a block file manager 162, a compression manager 154, and a non-transitory computer-readable medium. In such an example, the instructions, when executed by the processing device 199, configure the processing device 199 of the data platform 150 to perform various operations. For example, the instructions may configure the processing device 199 to determine an entropy value 186 for each of a plurality of data blocks 174 to obtain a corresponding plurality of entropy values ​​186. In an example, the instructions may configure the processing device 199 to reorganize the plurality of data blocks 174 by the block file manager 162 and based on the corresponding plurality of entropy values ​​186 to obtain a reorganized plurality of data blocks 174. Continuing with the example, the instructions may configure the processing device 199 to compress the reorganized plurality of data blocks 174 by the compression manager 154 to obtain a compressed block file 176. In this example of the data platform 150 , the instructions may configure the processing device 199 to store the compressed block file 176 via the storage system 105 , 115 in place of the plurality of data blocks 174 within the storage system 105 , 115 .

[0040] In another example, various aspects of the technology relate to a computer-readable storage medium having instructions that, when executed, configure a processing device 199 to perform various operations. In such an example, the instructions, when executed, may configure the processing device 199 to determine an entropy value 186 for each of a plurality of data blocks 174 to obtain a corresponding plurality of entropy values ​​186. In this example, the instructions, when executed, may configure the processing device 199 to reorganize the plurality of data blocks 174 based on the corresponding plurality of entropy values ​​186 to obtain a reorganized plurality of data blocks 174. Continuing with this example, the instructions, when executed, may configure the processing device 199 to compress the reorganized plurality of data blocks to obtain a compressed block file 176. In this example of a computer-readable storage medium, the instructions, when executed, may configure the processing device 199 to store the compressed block file 176 in place of the plurality of data blocks 174.

[0041] exist Figure 1In an example of data platform 150, data platform 150 includes compression manager 154, which provides compression services on behalf of data platform 150. Compression manager 154 can select and use one or more available compression algorithms 107. Compression manager 154 can obtain and / or evaluate one or more attributes 106 of data to be compressed by compression manager 154. In some examples, compression manager 154 organizes, compresses, and stores information generated by data platform 150. In some examples, compression manager 154 organizes, compresses, and stores information generated by one or more mobile devices 108 and one or more client devices 109. Compression manager 154 includes entropy calculator 158 for calculating entropy values ​​of data blocks 174. In some examples, compression manager 154 calculates bit value entropy value 186A. In some examples, compression manager 154 calculates byte value entropy value 186B. The calculated bit value entropy value 186A and byte value entropy value 186B are collectively referred to as "entropy value 186". In some examples, compression manager 154 selects compression algorithm 107 based on the calculated entropy value 186.

[0042] A plurality of pointers 173A, 173B, 173C, and 173D (collectively referred to as "pointers 173") point to and / or reference data blocks 174. In an operation to reorganize the plurality of data blocks 174 to obtain the reorganized plurality of data blocks 174, the set of data blocks 174 may be reorganized into a new order 175 by rearranging the pointers 173 linked to, pointing to, and / or referencing each of the individual data blocks 174. Reordering the pointers 173 pointing to the data blocks 174 rather than directly rearranging the data blocks 174 may save a significant amount of computational resources because there is no need to read and rewrite the data blocks to and from the storage systems 105, 115. For example, since the size of the pointers 173 is very small compared to the relatively large size of the data blocks 174, the amount of computation required to reorder and subsequently update and / or rewrite the pointers 173 pointing to the storage systems 105, 115 using the new order 175 may be significantly reduced.

[0043] As discussed above, when data is compressed into highly random block files, space utilization, compression efficiency, and computing resources (e.g., CPU cycles) may be wasted. The compression manager 154 can generate higher storage efficiency by applying preprocessing, although this will come at the expense of complexity and overhead of the data platform 150. Data files have varying degrees of randomness. For example, due to the high repetitiveness of the internal structure of structured files, such file formats tend to exhibit a high degree of order. In contrast, encrypted files and compressed files tend to exhibit a high degree of randomness or disorder. Encrypted files tend to become disordered because the encryption algorithm intentionally introduces disorder and complexity to the file for security measures. Compressed files tend to become disordered because the selected compression algorithm 107 replaces repeated bit sequences with syntax representing the original data content and structure in order to reduce storage space. The compression manager 154 can introduce an increasing order (e.g., reduce entropy) before compression by rearranging the data blocks 174. For example, the data blocks 174 can be organized in ascending or descending order, which places data blocks with similar entropy values ​​next to each other. In this manner, the reordered data blocks 174 may exhibit less entropy overall than the data blocks 174 prior to being reorganized. The lower overall entropy of the set of data blocks 174 used to create the compressed block file 176 may result in lower storage space consumption due to the higher compression efficiency achieved.

[0044] According to various aspects of the technology described in the present disclosure, the compression manager 154 can perform a series of operations to reorganize the data blocks 174 into a new order, thereby producing a higher compression efficiency. For example, the compression manager 154 can obtain and / or retrieve the data blocks 174 from the application system 102 and / or the file system 153. In this example, the compression manager 154 can receive pointers 173 pointing to all data blocks 174 stored locally, which are selected for creating a new block file 164. Continuing with this example, the compression manager 154 can reorder the pointers 173 and / or reorganize a certain identifier representing the data blocks 174 into a new order 175. In this example, the entropy calculator 158 of the compression manager 154 can calculate an entropy value 186 for each corresponding data block 174 by parsing the pointers 173 and / or references to each underlying data block 174. Continuing with the example, the compression manager 154 has the entropy value 186 calculated for each respective data block 174, sorts, reorders, rearranges, and / or resequences the pointers to the entropy value 186 calculated for each of the respective data blocks 174, thereby producing a new order 175. The compression manager 154 may sequentially apply compression to the data blocks 174 by dereferencing (e.g., following) the pointers 173 in the new order 175, sequentially moving through the data blocks 174 according to the new order 175 specified by the pointers to create a compressed block file 176. The compressed block file 176 is then stored, thereby reducing overall storage system consumption.

[0045] For example, when the compressed block file 176 consumes less storage resources than the data blocks 174 replaced by the compressed block file 176, the total storage consumption will be reduced because the compressed block file 176 replaces and / or replaces the corresponding data blocks 174. Since the compressed block file 176 is formed by the block files 174 reordered according to the calculated entropy value 186 of each of the corresponding data blocks 174, the storage space consumed by the compressed block file 176 can be reduced compared to a compressed block file formed by the data blocks 174 appearing in a random order.

[0046] The entropy of data represents the randomness and / or "disorder" inherent in such data (e.g., how disordered the data is). Entropy, based on concepts from thermodynamics and applied to information theory by Claude Shannon, provides a mechanism by which the randomness of data stored by a file system can be systematically measured. Shannon's theory defines a data communication system consisting of three elements: a data source, a communication channel, and a receiver. Shannon pointed out that the fundamental problem of communication is the ability of a receiver to identify data generated by a source based on the signal it receives over the channel. Shannon asserted that there is an absolute mathematical limit to the effectiveness of compressing data from a source onto a completely noise-free channel using lossless compression. Lossless compression is a class of data compression that allows the original data to be perfectly reconstructed from the compressed data without any loss of information. Lossless compression is possible because most real-world data exhibits statistical redundancy. In contrast, lossy compression only allows an approximation of the original data to be reconstructed, but typically with greatly improved compression ratios and, as a result, reduced file storage size.

[0047] Shannon defined entropy by the following formula:

[0048] H=-∑ i p i log(p i );where p i is the frequency of each symbol i (the sum), and if the logarithm base is 2, the result H is expressed in bits per symbol. Therefore, if the entropy value is close to 8, for example 7.98, the entropy value will mean that in each byte of data, 7.98 bits are essentially random. Encrypted and compressed files have an entropy value close to 8, indicating that the compressed file cannot be compressed further. In contrast, consider another example where there is a file filled with all zeros. The entropy value of such a file will be close to 0, indicating that the file is highly compressible.

[0049] In some examples, the processing device 199 may select a plurality of data blocks 174 from the storage systems 105, 115 to create a block file. In some examples, in response to selecting a plurality of data blocks 174 from the storage systems 105, 115 to create the block file 164, the processing device 199 determines an entropy value 186 for each of the selected plurality of data blocks 174. In some examples, the processing device 199 may determine each entropy value 186 by calculating the entropy value 186 of each data block 174 according to the formula H = -∑ i p i log(p i ). In some examples, the term H represents the entropy value 186 as calculated by the processing device 199. In some examples, the term i represents the index of each of the plurality of symbols. In some examples, the term p iis the frequency of each of the plurality of symbols i. In some examples, when the logarithm base is 2, the entropy value 186 represented by the term H is represented in terms of the number of bits per symbol.

[0050] In some examples, the processing device 199 may select a compression algorithm 107 based on the calculated entropy value 186 as determined by the entropy calculator 158. In some examples, the processing device 199 may select a compression algorithm 107 from a plurality of compression algorithms 107 based on properties of the plurality of data blocks 174. For example, the compression manager 154 may select a compression algorithm 107 based on properties of the data blocks 174 as obtained by the compression manager 154. In some examples, the compression manager 154 may compress the block file 174 using the selected compression algorithm 107. The processing device may apply any of a plurality of compression algorithms to generate and / or create a compressed block file. In some examples, the processing device compresses the plurality of data blocks 174 using the new order 175 to generate a compressed block file 176.

[0051] In some examples, the compression algorithm may be selected based on the determinable properties of the plurality of data blocks. In some examples, the compression algorithm 107 may be selected based on the properties of the block file 164 created from the plurality of data blocks 174 before the block file 164 is compressed into the compressed block file 176. In some examples, the compression algorithm 107 may be selected based on the properties of the underlying files embodied in the plurality of data blocks 174 stored to the storage system. For example, music files stored in an analog waveform format may preferably utilize a different compression algorithm 107 than digitized music. Video files may preferably utilize a different compression algorithm 107 than database backup files. Data processing files (e.g., *.doc, *.docx, *.xls, googledocs, *.txt, etc.) may preferably utilize a different compression algorithm 107 than *.pdf files and image files. In some examples, the properties of the plurality of data blocks 174 may be determined after the plurality of data blocks 174 within the storage system are reorganized and / or reordered using the new order 175. In some examples, attributes of plurality of data chunks 174 may be determined prior to reorganizing and / or reordering the plurality of data chunks.

[0052] The storage system 115 includes one or more storage devices 140A to 140X (collectively referred to as "storage devices 140"). The storage device 140 may represent one or more physical or virtual computers and / or storage devices that include or otherwise have access to storage media. Such storage media may include one or more of a flash drive, a solid state drive (SSD), a hard disk drive (HDD), an optical disk, an electrically programmable memory (EPROM) or an electrically erasable programmable (EEPROM) memory form and / or other types of storage media. Different storage devices of the storage device 140 may have different types of storage media combinations. Each of the storage devices 140 may include system memory. Each of the storage devices 140 may be a storage server, a network attached storage (NAS) device, or may represent disk storage of a computer device. The storage system 115 may include a redundant array of independent disks (RAID) system. The storage system 115 may be capable of storing much larger amounts of data than the storage system 105. The storage device 140 may also be configured for long-term information storage that is more suitable for archival purposes.

[0053] In some examples, storage systems 105 and / or 115 may be storage systems deployed and managed by a cloud storage provider and are referred to as "cloud storage systems". Example cloud storage providers include, for example, AMAZON WEBSERVICES (AWS TM ), MICROSOFT,INC. DROPBOX,INC.'s DROPBOX TM 、ORACLE CLOUD of ORACLE,INC. TM AND GOOGLE,INC.’S GOOGLE CLOUD PLATFORM TM (GCP). In some examples, storage system 115 is located with storage system 105 in a data center, on-premises, or in a private cloud, public cloud, or hybrid private / public cloud. Storage system 115 can be considered a "backup" or "secondary" storage system for primary storage system 105. Storage system 115 can be referred to as an "external target" for archive 142. When deployed and managed by a cloud storage provider, storage system 115 can be referred to as "cloud storage."

[0054] The storage system 115 may include one or more interfaces for managing data transfer between the storage system 105 and the storage system 115 and / or between the application system 102 and the storage system 115. The data platform 150 supporting the application system 102 may rely on the primary storage system 105 to support latency-sensitive applications. However, because the storage system 105 is generally more difficult to scale or more expensive, the data platform 150 may use the secondary storage system 115 to support secondary use cases, such as backup and archiving. In general, a file system backup is a copy of the file system 153 that is used to support protection of the file system 153 for rapid recovery, which is usually due to the loss of some data in the file system 153, and a file system archive ("archive") is a copy of the file system 153 that is used to support long-term retention and review. A "copy" of the file system 153 may include such data as is needed to restore or view the state of the file system 153 at the time of the backup or archive.

[0055] The compression manager 154 may archive the file system data of the file system 153 at any time according to an archive policy, which specifies, for example, archive periodicity and timing (daily, weekly, etc.), the file system data to be archived, archive retention period, storage location, access control, etc. The initial archive of the file system data may correspond to the state of the file system data at the initial archive time (the archive creation time of the initial archive). Depending on the archive policy, the initial archive may include a complete archive of the file system data, or may include a partially complete archive of the file system data. For example, the initial archive may include all objects of the file system 153 or one or more selected objects of the file system 153, including data blocks 174 in an ordered state or an unordered state.

[0056] One or more subsequent incremental archives of the file system 153 may correspond to respective states of the file system 153 at respective subsequent archive creation times (i.e., after the archive creation time corresponding to the initial archive). The subsequent archives may include incremental archives of the file system 153. The subsequent archives may correspond to incremental archives of one or more objects of the file system 153, including data blocks 174 in an ordered state or an unordered state. Some of the file system data of the file system 153 stored on the storage system 105 at the initial archive creation time may also be stored on the storage system 105 at the subsequent archive creation time. The subsequent incremental archives may include data that was not previously archived to the storage system 115. The compression manager 154 may deduplicate the file system data included in the subsequent archives based on the file system data included in one or more previous archives (including the initial archive) to reduce the amount of storage used. (The "time" referred to in the present disclosure may refer to a date and / or a time. A time may be associated with a date. For example, multiple archives may appear at different times on the same date.)

[0057] In system 100, compression manager 154 may coordinate the sorting, compression, and storage of information to one or more data stores. In some examples, compression manager 154 may reorder data blocks 174 and compress data blocks 174 into block files 164. In some examples, compression manager 154 may operate in coordination with block file manager 162 to sort, compress, and store data blocks 174 into block files 164. In some examples, entropy calculator 158 may calculate an entropy value 186 for each of a plurality of data blocks 174. In some examples, block file manager 162 may reorder a plurality of data blocks 174 in ascending or descending order according to their corresponding entropy values ​​186. Figure 1 In the example of , data blocks 174A, 174B, 174C, and 174D are reorganized into a new order 175 according to their corresponding entropy values ​​186, thereby producing an order of data blocks 174C, 174A, 174D, and 174B. In some examples, block file manager 162 compresses the reordered data blocks 174C, 174A, 174D, and 174B into block file 164. Figure 1 In the example of , compressed block file 176 is created by block file manager 162, which compresses reordered data blocks 174C, 174A, 174D, and 174B into a single compressed block file 176 using new order 175 of data blocks 174. Compressed block file 176 may be stored by storage system 105 using storage device 180, and / or stored within archive 142 of storage system 115. In some examples, block file 164 is written to archive 142. In some examples, block file 164 replaces data block 174. In some examples, block file 164 replaces data block 174 by updating references and metadata pointing to data block 174 to point to block file 164 instead.

[0058] In some examples, the processing device 199 may use the entropy calculator 158 to calculate an entropy value 186 for each of the plurality of data blocks 174. In some examples, in response to determining the entropy value 186 for each of the plurality of data blocks 174, the processing device 199 organizes the plurality of data blocks 174 into a new order 175 according to the entropy value 186 for each of the plurality of data blocks 174. In some examples, the block file manager 162 writes the plurality of data blocks 174 to the storage system 105, 115. In some examples, pointers 173 that reference the plurality of data blocks 174 are written to the storage system 105, 115 using the new order 175 into which the plurality of data blocks 174 are reorganized and / or reordered. In some examples, the compression manager 154 generates a compressed block file 176 by compressing the plurality of data blocks 174 using the new order 175 of the pointers 173 that reference the plurality of data blocks 174. In some examples, the compression manager 154 replaces and / or substitutes the plurality of data blocks 174 on the storage system 105, 115 with the compressed block file 176. In other words, the compressed block file 176 written to the storage system 105, 115 replaces the previous variant of the block file 164 in an uncompressed format and replaces the previous variant of the data block 174 now embodied within the compressed block file 176. In some examples, the previous variant of the block file 164 may be overwritten. In some examples, the previous variant of the block file 164 may be replaced by updating a pointer within the storage system 105, 115 that references the previous variant of the block file 164 to instead reference the newly compressed block file 176. In such an example, replacing the previous variant of the data with the compressed block file 176 frees up storage system space and may provide a more efficient data storage environment within the data platform 150.

[0059] In some examples, the processing device may rewrite and / or update the pointers 173 pointing to the plurality of data blocks 174 of the storage system 105, 115 in ascending or descending order using the new order 175 determined by the block file manager 162. In some examples, each of the pointers 173 that reference each of the plurality of data blocks stored by the storage system 105, 115 may be stored as a node within a linked list. In some examples, the processing device may sequentially organize the nodes within the linked list corresponding to the pointers 173 that reference each of the plurality of data blocks 174 stored by the storage system 105, 115 into one of a descending order or an ascending order based on establishing the new order 175 for the plurality of data blocks 174. In this manner, a sequential arrangement of the plurality of data blocks 174 may be established without incurring the computational burden of relocating the plurality of data blocks 174 on a physical medium.

[0060] Figure 2is a block diagram illustrating another example system 200 that provides a more efficient data storage environment with data deduplication, encryption, and entropy thresholds, in accordance with one or more techniques of this disclosure. Figure 2 The system 200 can be described as Figure 1 Examples or alternative implementations of the system 100. Figure 1 Described in the background Figure 2 one or more aspects of. Figure 2 In the example of , system 200 includes network 111, data platform 150, and storage system 115. Figure 2 In the example of FIG. 1 , the network 111, the data platform 150, and the storage system 115 may correspond to Figure 1 network 111, data platform 150 and storage system 115. Figure 2 The data platform 150 and storage system 115 can apply the technology according to the present disclosure, including Figure 1 The network 113 provides application services to one or more mobile devices 108 and one or more client devices 109. Different instances of the data platform 150 and / or storage system 115 may be deployed by different cloud storage providers, the same cloud storage provider, an enterprise, or other entities.

[0061] exist Figure 2 In an example of FIG. 1 , data platform 150 includes encryption manager 245 having an encrypted compressed block file 286 created by encryption manager 245. In some examples, encrypted compressed block file 286 is created by encryption manager 245 and stored via storage system 105 and / or storage system 115. Within data platform 150, compression manager 154 may include entropy calculator 158 to calculate entropy value 186. Compression manager 154 may include a configurable specified entropy value threshold 287. In some examples, entropy value threshold 287 provides a configurable value at which compression manager 154 may evaluate whether to compress data blocks 174, data chunks, and / or files stored by storage systems 105, 115.

[0062] exist Figure 2In an example of , storage system 115 includes archive 142, which may store compressed block files 176 and / or encrypted compressed block files 286. Storage system 115 may include block file manager 162 configured with deduplicator 240. In some examples, deduplicator 240 may remove identical or other unwanted copies of files, data chunks, and / or data chunks 174. In some examples, deduplicator 240 may operate on a set of data chunks 241 to create a deduplicated set of data chunks 242. In some examples, block file manager 162 may select data chunk 174 from the set of data chunks 241. In some examples, block file manager 162 selects data chunk 174 from the deduplicated set of data chunks 242.

[0063] In some examples, the processing device may de-duplicate the set of data chunks 241 stored by the storage systems 105, 115 via the de-duplicator 240 to create a de-duplicated set of data chunks 242. In some examples, the processing device 199 of the data platform 150 may select a plurality of data chunks 174 from the de-duplicated set of data chunks 174 to be included in the chunk file 164.

[0064] Deduplication is a process that eliminates excess copies of data and significantly reduces storage capacity requirements. Deduplication can be run as an inline process when data is written to the storage system 105, 115, and / or as a background process to eliminate duplicate data after the data is written to disk or otherwise stored by the storage system 105, 115. In some examples, deduplication can be run on the data file before the data block 174 is created. In some examples, deduplication can be run on the data block 174 to eliminate one or more identical copies of the data block 174, thereby generating a deduplicated set of data blocks 242. In some examples, deduplication can be applied to a first data block 174 identified as having one or more identical copies by defining the one or more identical copies of the first data block 174 as a reference and / or pointer to the first data block. In this example, the data within the first data block 174 can be retained in an unmodified form and one or more identical copies can be replaced by a reference and / or pointer to the first data block 174, thereby significantly reducing the space consumed by the first data block 174 and one or more identical copies of the first data block 174 on the storage systems 105, 115.

[0065] For example, a job seeker submits his resume to multiple job postings on a job search platform. Each of the resumes may be identical, but has been submitted multiple times. As an illustrative example, a copy of multiple identical resumes is retained by the job search platform, and one or more identical copies of the resume are redefined as a reference or pointer to a retained copy of the resume. In another example, consider two users of a music streaming platform, each of which downloads a song to their personal library, which is stored in the cloud by the music streaming platform. Similar to the resume example, only one copy of the song needs to be retained, and the second copy of the song of the second user is redefined as a pointer and / or reference to the first copy of the song retained by the music streaming platform, thereby significantly reducing the space consumption of the underlying storage system 105, 115.

[0066] In some examples, the processing device may generate an encrypted compressed block file 286 via the encryption manager 245. In some examples, the processing device 199 may encrypt the compressed block file 176 into a single file to generate the encrypted compressed block file 286. Encrypted data tends to have a very high entropy value, and therefore, it may be preferable to compress the reorganized data blocks 174 into the compressed block file 176 before applying encryption.

[0067] In some examples, the entire compressed block file 176 may be encrypted, and the non-encrypted version of the compressed block file 176 may be replaced within the storage system with an encrypted compressed block file 286 variation of the compressed block file 176. However, encryption may optionally be performed before the block file is created and / or before the block file is compressed. In some examples, some or all of the files that make up each of the plurality of data blocks 174 are encrypted. In some examples, the plurality of data blocks 174 are each encrypted before the compressed block file is created. In some examples, a block file is created from the plurality of data blocks 174, and the block file may be encrypted before being compressed. In some examples, the compressed block file is created by compressing the plurality of data blocks 174 into a single block file, and the processing device may encrypt the compressed block file to generate an encrypted compressed block file.

[0068] In some examples, the processing device may reorder the plurality of data blocks 174 into ascending or descending order based on the entropy values ​​186 of the plurality of data blocks 174. In some examples, in response to reordering the plurality of data blocks 174 into ascending or descending order, the processing device 199 may rewrite and / or update the pointers 173 pointing to the plurality of data blocks 174 of the storage systems 105, 115 in ascending or descending order.

[0069] In some examples, the plurality of data blocks 174 may be organized, reorganized, and / or reordered into descending or ascending order to improve compression efficiency. For example, experiments have shown that when the plurality of data blocks 174 are reordered into descending or ascending order (when written to a physical storage medium or referenced by a pointer 173 pointing to the plurality of data blocks 174), efficiency gains of more than 10% have been achieved using one or more of the techniques described herein. In some examples, the processing device may access the pointers in the new order 175 to retrieve the plurality of data blocks 174 and compress the plurality of data blocks 174 into a compressed block file 176. In some examples, the plurality of data blocks 174 may be loaded into a local memory for the compression algorithm 107 using the new order 175 of the pointers 173, and the plurality of data blocks 174 may be processed online and / or sequentially. In other words, the compression algorithm 107 may process the plurality of data blocks 174 in the order in which they are loaded into memory, which corresponds to the new order 175 reorganized into by the pointers 173. In some examples, the pointers 173 may be maintained within a linked list, and the plurality of data blocks 174 may be reorganized by updating the order of the pointers 173 pointing to the plurality of data blocks 174 within the linked list. The linked list may provide a linear collection of data elements having an order that is not based on the physical layout of the data elements within the storage system 105, 115, but rather, the linked list may be organized as a data structure consisting of a collection of nodes, each node having a pointer 173 pointing to one of the plurality of data blocks 174, wherein the collection of nodes represents a sequence. Thus, the nodes may be reorganized within the linked list, changing the order of the sequence without relocating the physical layout of any of the plurality of data blocks 174. In some examples, the pointers 173 may be maintained within the file system 153, and the plurality of data blocks 174 may be reorganized by updating the order of the pointers 173 pointing to the plurality of data blocks 174 as maintained by the file system 153.

[0070] The compression efficiency gain exhibits a generally inverse correlation with the calculated entropy value 186 for each compression unit, whether the compression unit is a file, a data block, a data block 174, or a block file 164. In some examples, the entropy value 186 may be calculated in a range of 1 to 8. In other examples, the entropy value 186 may be calculated as a value between 0 and 1. Regardless of the range or scale used, higher entropy values ​​186 are generally associated with a higher degree of disorder and / or randomness, and therefore produce lower compression efficiency.

[0071] Conversely, a lower entropy value 186 is generally associated with a lower degree of disorder and / or randomness, and therefore, the compression manager 154 may achieve a higher compression efficiency. For example, an entropy value of "2" in the range of 1 to 8 or an entropy value of "0.2" in the range of 0 to 1 may indicate low entropy (e.g., a low degree of disorder and / or randomness), and therefore, the selected compression algorithm 107 may achieve a higher compression efficiency by exploiting the internal structure within the low entropy data block 174. However, an entropy value of "8" in the range of 1 to 8 or an entropy value of "0.99" in the range of 0 to 1 may indicate excessively high entropy (e.g., a high degree of disorder and / or randomness), and therefore, the selected compression algorithm 107 may produce little or no compression efficiency due to the randomness and lack of structure within the high entropy data block 174. Counterintuitively, a compression algorithm applied to a high entropy data block 174 may result in a "compressed" data block that is larger in size than a corresponding uncompressed variant of the same data block 174. In other words, due to the high degree of disorder and / or randomness, the compression algorithm 107 may increase the size of the data block 174 when storing the data block in a "compressed" form. The reason for this is that more space is required to store the syntax and data describing the replacement elements within the data block 174 when stored in a compressed format rather than simply storing the same data in a raw and uncompressed form (without any compression syntax).

[0072] Thus, in some examples, the data blocks 174 may be evaluated according to the specified and configurable entropy value threshold 287 before the compression algorithm 107 is applied to the data blocks 174. For example, a data block 174 having a calculated entropy value 186 of "6" may exceed the configurable specified entropy value threshold 287 of "5" for the entropy value 186. In such an example, the processing device 199 may affirmatively filter the data block 174 out of the block file 164. In some examples, the processing device 199 may eliminate the data blocks 174 whose calculated entropy values ​​186 do not meet the entropy value threshold 287 from the plurality of data blocks 174 to be compressed into the compressed block file 176. In some examples, the plurality of data blocks 174 may be selected as a subset or portion of the set of data blocks 241 based on the selected subset of the data blocks 174 having a calculated entropy value 186 less than the entropy value threshold 287.

[0073] In some examples, the pointers 173 that reference each of the plurality of data blocks 174 may be sequentially arranged in ascending or descending order using a linked list. In this manner, the sequential arrangement of the plurality of data blocks 174 may be processed using a new order 175 established for the plurality of data blocks 174 based on the entropy values ​​corresponding to each of the plurality of data blocks. In other words, data blocks 174 having similar entropy values ​​may be positioned adjacent to each other by reorganizing the pointers 173 using the new order 175. When run-length encoding (RLE) is used, more efficient compression may be obtained after reordering the plurality of blocks into descending or ascending order by increasing the length and number of runs of single-valued data spanning one or more of the plurality of data blocks 174. Run-length encoding is a lossless compression technique in which a sequence embodying redundant data is stored as a single data value that represents a repeated block of redundant data and the number of times the redundant data occurs in an underlying data block 174 or within a data sequence spanning the plurality of data blocks 174. In some examples, during a subsequent decoding and / or decompression stage, the run-length encoding information may be used to accurately reconstruct the original uncompressed data of the data block 174. Reordering the plurality of data blocks 174 into descending or ascending order according to the corresponding entropy value 186 calculated for each of the plurality of data blocks 174 may produce higher compression efficiency by reducing the overall disorder and / or randomness across the span of the newly organized data blocks 174. In other words, the data blocks 174 sequentially organized within the storage system 105, 115 according to the entropy values ​​186 of the data blocks 174 may reduce the disorder and / or randomness of the plurality of data blocks 174 that make up the block file 164. Higher efficiency compression may result from the adjacent placement of similar data structures, the adjacent placement of similar data sequences, and / or the adjacent placement of similar file formats.

[0074] In some examples, the processing device calculates an entropy value 186 for each data block 174 within the set of data blocks 241 stored by the storage system 105, 115. In some examples, the processing device compares each data block 174 within the set of data blocks 241 to an entropy value threshold 287, possibly via the block file manager 162. In some examples, in response to comparing each data block 174 within the set of data blocks 241 to the entropy value threshold 287, the processing device 199 selects a plurality of data blocks 174 from the set of data blocks 241 based on the selection of each data block 174 as a plurality of entropy values ​​186 that satisfy the entropy value threshold 287. In other words, a portion or subset of the data blocks 174 in the set of data blocks 241 that satisfy the entropy value threshold 287 may be selected to be included in the block file 164, included in the compressed block file 176, and / or included in the cryptographically compressed block file 286.

[0075] In some examples, the entropy value threshold 287 may be configured based on a tradeoff between the computational cost and resources required to compress the data blocks 174 and / or block files 164 and the compression efficiency gain resulting from performing any of the one or more techniques described herein. For example, for an entropy value range of 1 to 8, a configurable entropy value threshold 287 of "2" may ensure that all selected data blocks 174, once compressed, will exhibit high efficiency compression gain, thus providing sufficient return when measured in terms of storage system 105, 115 space savings compared to the computational resources required to achieve such storage system 105, 115 space savings. However, this low entropy value threshold of "2" may result in the desired potential storage system 105, 115 space savings not being achieved. In a variable computing demand environment, such as a data platform 150 that experiences higher computing loads during certain periodic cycles (e.g., nightly, weekly, end of quarter, end of year, etc.), the entropy threshold 287 may be dynamically configured so that more possible storage system 105, 115 space savings are achieved during periods of low computing demand by consuming excess or otherwise unused computing capacity of the data platform 150. Similarly, compression requirements may be classified, qualified, demoted, or otherwise configured as "backend" or "overhead" computing loads that are configured to be processed during periods of low computing demand for the data platform 150, thereby allowing higher computing efficiency to be achieved by accepting higher computing costs despite periods of low computing demand.

[0076] Figure 3 is a block diagram illustrating an example system 300 in accordance with techniques of this disclosure. Figure 3 The system 300 can be described as Figure 1 System 100 or Figure 2 Examples or alternative implementations of the system 200. Figure 1 and Figure 2 Description in the background Figure 3 one or more aspects of .

[0077] exist Figure 3 In the example of , system 300 includes network 111, data platform 150 implemented by computing system 302, and storage system 115. Figure 3 In the example, the network 111, the data platform 150 and the storage system 115 may correspond to Figure 1 and Figure 2 1 , a data platform 150, and a storage system 115. Although only one archival storage system 115 is depicted, the data platform 150 may also apply techniques according to the present disclosure using multiple instances of the archival storage system 115. Different instances of the storage system 115 may be deployed by different cloud storage providers, the same cloud storage provider, an enterprise, or other entities.

[0078] The computing system 302 may be implemented as any suitable computing system, such as one or more server computers, workstations, mainframes, appliances, cloud computing systems, and / or other computing systems that may be capable of performing operations and / or functions described according to one or more aspects of the present disclosure. In some examples, the computing system 302 represents a cloud computing system, a server farm, and / or a server cluster (or a portion thereof) that provides services to other devices or systems. In other examples, the computing system 302 may represent or be implemented by one or more virtualized computing instances (e.g., virtual machines, containers) of a cloud computing system, a server farm, a data center, and / or a server cluster.

[0079] exist Figure 3 In the example of , the computing system 302 may include one or more communication units 315, one or more input devices 317, one or more output devices 318, and one or more storage devices of the local storage system 105. The storage system 105 includes an interface module 326, a file system manager 152, a compression manager 154, an entropy calculator 158, and a compression algorithm 107. The storage system 105 may create or generate information during operation, so that the storage system 105 may also include data blocks 174, a new order 174 of the data blocks 174, an entropy value 186 calculated for the data blocks 174, and / or one or more compressed block files 176. The storage system 105 may optionally store the block file 164 and the encrypted compressed block file. One or more of the devices, modules, storage areas, or other components of the computing system 302 may be interconnected to enable communication between the components (physically, communicatively, and / or operationally). In some examples, such connectivity may be provided through a communication channel (e.g., communication channel 312), which may represent one or more of a system bus, a network connection, an inter-process communication data structure, or any other method for transferring data.

[0080] The computing system 302 includes a processing device. Figure 3 In the example of FIG. 1 , the processing device includes one or more processors 313, wherein the one or more processors are configured to implement the functions associated with or related to the computing system 302. Figure 3and described below and may perform the functionality associated with one or more modules shown in and described below and / or execute instructions associated with the above. One or more processors 313 may be a processing circuit system that performs operations according to one or more aspects of the present disclosure, may be a part of the processing circuit system, and / or may include the processing circuit system. Examples of processor 313 include a microprocessor, an application processor, a display controller, an auxiliary processor, one or more sensor hubs, and any other hardware configured to function as a processor, processing unit, or processing device. The computing system 302 may use one or more processors 313 to perform operations according to one or more aspects of the present disclosure using software, hardware, firmware, or a mixture of hardware, software, and firmware resident in and / or executed at the computing system 302.

[0081] One or more communication units 315 of computing system 302 may communicate with devices external to computing system 302 by transmitting and / or receiving data, and in some aspects may function as both an input device and an output device. In some examples, communication unit 315 may communicate with other devices over a network. In other examples, communication unit 315 may send and / or receive radio signals over a radio network (such as a cellular radio network). In other examples, communication unit 315 of computing system 302 may transmit and / or receive satellite signals over a satellite network. Examples of communication unit 315 include a network interface card (e.g., an Ethernet card), an optical transceiver, a radio frequency transceiver, a GPS receiver, or any other type of device that can send and / or receive information. Other examples of communication unit 315 may include a device that is capable of transmitting information over a wireless network. GPS, NFC, and cellular networks (e.g., 3G, 4G, 5G) and mobile devices devices that communicate with other devices such as radios and Universal Serial Bus (USB) controllers. Such communications may comply with, implement, or follow appropriate protocols, including Transmission Control Protocol / Internet Protocol (TCP / IP), Ethernet, NFC or other technologies or protocols.

[0082] One or more input devices 317 may represent any input device of computing system 302 not otherwise separately described herein. Input device 317 may obtain, generate, receive, and / or process input. For example, one or more input devices 317 may generate or receive input from a network, a user input device, or any other type of device for detecting input from a person or a machine.

[0083] One or more output devices 318 may represent any output device of computing system 302 that is not otherwise described separately herein. Output device 318 may generate, present, and / or process output. For example, one or more output devices 318 may generate, present, and / or process output in any form. Output device 318 may include one or more USB interfaces, video and / or audio output interfaces, or any other type of device capable of generating tactile, audio, visual, video, electrical, or other output. Some devices may be used as both input devices and output devices. For example, a communication device may send data to other systems or devices via a network, and may receive data from other systems or devices via a network.

[0084] One or more storage devices of the local storage system 105 within the computing system 302 may store information for processing during operation of the computing system 302, such as random access memory (RAM), flash memory, solid state disk (SSD), hard disk drive (HDD), etc. The storage device may store program instructions and / or data associated with one or more of the modules described according to one or more aspects of the present disclosure. One or more processors 313 and one or more storage devices may provide an operating environment or platform for such modules, which may be implemented as software, but in some examples may include any combination of hardware, firmware, and software. One or more processors 313 may execute instructions, and one or more storage devices of the storage system 105 may store instructions and / or data of one or more modules. The combination of the processor 313 and the local storage system 105 may retrieve, store, and / or execute instructions and / or data of one or more applications, modules, or software. The processor 313 and / or storage device of the local storage system 105 may also be operably connected to one or more other software and / or hardware components, including but not limited to one or more of the components of the computing system 302 and / or one or more devices or systems shown as connected to the computing system 302.

[0085] The file system manager 152 may perform functions related to providing the file system 153, as described above with respect to Figure 1 and Figure 2The file system manager 152 may generate and manage file system metadata 332 for constructing file system data 330 of the file system 153, and store the file system metadata 332 and the file system data 330 to the local storage system 105. The file system metadata 332 may include one or more trees, which describe objects within the file system 153 and the file system 153 hierarchy, and may be used to write or retrieve objects within the file system 153. The file system metadata 332 may indirectly reference the data block 174 by using the pointer 173, and then retrieve the referenced data block 174 according to the pointer. The compression manager 154 may reference the file system metadata 332 to support the execution of block file operations and management. The file system manager 152 may interact and / or operate in coordination with one or more modules of the computing system 302, including the interface module 326 and the compression manager 154.

[0086] The compression manager 154 may perform compression functions related to: chunking files and data into data chunks 174, calculating entropy values ​​186 via the entropy calculator 158; and organizing the data chunks 174 into a new order 175; compressing selected data chunks 174 into compressed chunk files 176; and encrypting data chunks 174, encrypting chunk files 164, and / or encrypting compressed chunk files 176, as described above with respect to Figure 1 and Figure 2 Described, including the operations described above with respect to coordination with block file manager 162 .

[0087] The interface module 326 may implement an interface through which other systems or devices may determine the operation of the file system manager 152, the compression manager 154, and / or the block file manager 162. Another system or device may communicate via the interface of the interface module 326 to specify one or more entropy thresholds 287.

[0088] System 300 may be modified to implement Figure 1 System 100 or Figure 2 In some examples of the modified system 300, the storage system 105 and / or 115 may use the encryption manager 245 to perform encryption operations. In some examples of the modified system 300, the storage system 105 and / or 115 may use the deduplication manager 240 to perform deduplication operations. In some examples of the modified system 300, the storage system 105 and / or 115 includes both the compression manager 154 and the block file manager 162 to perform the above-referenced Figure 1 and / or Figure 2 One or more techniques described.

[0089] In some examples, the data platform 150 includes a processing device (e.g., a processor 313), a storage system 105, 115, a block file manager 162, a compression manager 154, and a non-transitory computer-readable medium. In some examples, the instructions, when executed by the processing device, configure the processing device to perform operations. In some examples, in response to determining an entropy value 186 for each of the plurality of data blocks 174, the processing device organizes the plurality of data blocks 174 into a new order 175. In some examples, the plurality of data blocks 174 are organized according to the entropy value 186 for each of the plurality of data blocks 174. For example, the plurality of data blocks 174 may be organized or reorganized into an ascending order or a descending order. In some examples, the processing device writes and / or updates a pointer 173 to the plurality of data blocks 174 of the storage system 105, 115 using the new order 175, the plurality of data blocks 174 being organized into the new order. In some examples, the processing device generates a compressed block file using the compression manager 154. In some examples, the processing device compresses the plurality of data blocks 174 using the new order 175. In some examples, the compression manager replaces the plurality of data blocks 174 on the storage system 105, 115 with the compressed block file 176.

[0090] Although the techniques described in this disclosure are primarily described with respect to archive functions performed by the compression manager 154 and the block file manager 162 of the data platform 150, similar techniques may also be applied additionally or alternatively to backup, replica, clone, or snapshot functions performed by the data platform 150. In such cases, the compression manager 154 and the block file manager 162 may operate on backup, replica, clone, snapshot, or other data archived by, stored within, or accessible to the data platform 150.

[0091] Figure 4 is a flow chart illustrating an example mode of operation of a computing device to create a more efficient block file by using an entropy metric in accordance with the techniques of this disclosure. Figure 1 System 100, Figure 2 System 200 and Figure 3 The computing device 302 and storage systems 105, 115 describe the operating mode.

[0092] The data platform 150 may select data blocks for the block file (405). For example, the processing device 199 of the data platform 150 may select a plurality of data blocks 174 from a set of data blocks. The data platform 150 may calculate an entropy value for each data block (410). In some examples, the processing device determines an entropy value for each data block to be included in the block file being created.

[0093] The data platform 150 may reorganize the data blocks according to the entropy value (415). For example, in response to determining the entropy value of each of the plurality of data blocks, the processing device may reorganize the plurality of data blocks into a new order according to the entropy value calculated for each of the plurality of data blocks to obtain a reorganized plurality of data blocks. The data platform 150 may compress the reorganized data blocks to obtain a compressed block file (420). The data platform 150 may store the compressed block file to replace the data blocks (425). For example, the processing device may store the compressed block file in a storage system to replace the plurality of data blocks used to create the compressed block file. In some examples, the processing device may replace the plurality of data blocks used to create the compressed block file with the compressed block file. In some examples, the processing device may update the pointers to the plurality of data blocks used to create the compressed block file with one or more pointers to the compressed block file.

[0094] In some examples, the processing device may deduplicate a set of data blocks stored by the storage system to create a deduplicated set of data blocks. In some examples, the processing device may select a plurality of data blocks from the deduplicated set of data blocks for creating a compressed block file. In some examples, the processing circuitry may generate an encrypted compressed block file. In some examples, the processing circuitry may encrypt the compressed block file as a single file to generate an encrypted compressed block file.

[0095] In some examples, the processing device may calculate an entropy value for each data block within a set of data blocks stored by the storage system. The processing device of the data platform may use a block file manager to compare each data block within the set of data blocks with an entropy value threshold. In response to comparing each data block within the set of data blocks with the entropy value threshold, the processing device may select multiple data blocks from the set of data blocks based on the entropy value of each data block selected to satisfy the entropy value threshold.

[0096] Encryption may be applied to the compressed block file to improve the security of information stored to the storage system. In some examples, the processing device may encrypt the compressed block file as a single file to obtain an encrypted compressed block file. Encrypting the block file after compression may be preferred because compressing the previously encrypted block file may produce little or no compression efficiency gain due to the high entropy (e.g., high disorder) of the encrypted data. In some examples, the processing device may select a compression algorithm from a plurality of compression algorithms based on properties of the plurality of data blocks. In some examples, the processing device may compress the reorganized plurality of data blocks using the selected compression algorithm to obtain a compressed block file.

[0097] In some examples, the processing device may reorganize the pointers that reference each of the plurality of data blocks stored by the storage system into one of descending or ascending order according to the entropy value of each of the plurality of data blocks. In some examples, in response to reorganizing the pointers that reference each of the plurality of data blocks into ascending or descending order, the processing device may update the pointers that reference each of the plurality of data blocks within the storage system into ascending or descending order. In some examples, the pointers that reference each of the plurality of data blocks stored by the storage system may be stored as nodes within a linked list. In some examples, the processing device may sequentially organize the nodes within the linked list corresponding to the pointers that reference each of the plurality of data blocks stored by the storage system into one of descending or ascending order according to establishing a new order for the plurality of data blocks. In this way, an ascending or descending sequence of data blocks may be established based on the entropy value obtained without relocating the data blocks within the storage system by reorganizing the pointers and / or nodes that reference the data blocks.

[0098] In some examples, the processing device may select multiple data blocks from the data storage area for creating a block file. In some examples, in response to selecting multiple data blocks from the data storage area for creating a block file, the processing device may determine an entropy value for each of the selected multiple data blocks. For example, the processing device may calculate the entropy value or otherwise obtain the entropy value. In some examples, the processing device may calculate the entropy value using an entropy formula. In some examples, the processing device calculates each entropy value according to the following formula: H = -∑ i p i log(p i ). In some examples, the term H represents the entropy value 186 as calculated by the processing device 199. In some examples, the term i represents the index of each of the plurality of symbols. In some examples, the term p i is the frequency of each of the plurality of symbols i. In some examples, when the logarithm base is 2, the entropy value 186 represented by the term H is represented in terms of the number of bits per symbol.

[0099] For the processes, devices, and other examples or illustrations described herein, including in any flowchart or flow diagram, certain operations, actions, steps, or events included in any of the techniques described herein may be performed in a different order, may be added, merged, or omitted entirely (e.g., not all described actions or events are necessary for the practice of the technique). In addition, in some examples, operations, actions, steps, or events may be performed simultaneously, such as by multithreading, interrupt processing, or multiple processors, rather than sequentially. In addition, certain operations, actions, steps, or events may also be automatically performed even if not explicitly identified as being automatically performed. Moreover, certain operations, actions, steps, or events described as being automatically performed may alternatively not be automatically performed, but, in some examples, such operations, actions, steps, or events may be performed in response to an input or another event.

[0100] The detailed descriptions set forth herein in conjunction with the accompanying drawings are intended as descriptions of various configurations and are not intended to represent the only configurations in which the concepts described herein may be practiced. The detailed descriptions include specific details for the purpose of providing a thorough understanding of the various concepts. However, it is apparent to those skilled in the art that these concepts may be practiced without these specific details. In some instances, in order to avoid confusion of such concepts, well-known structures and components are shown in block diagram form.

[0101] According to one or more aspects of the present disclosure, when the context does not dictate otherwise, the term "or" may be interrupted to "and / or". In addition, although phrases such as "one or more" or "at least one" may have been used in some instances, they are not used in other instances; if the context does not dictate otherwise, these instances may be interpreted as having an implicit meaning in the absence of such language.

[0102] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on and / or transmitted via a computer-readable medium as one or more instructions or codes and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium, such as a data storage medium, or a communication medium, including any medium that facilitates the transfer of a computer program from one place to another (e.g., according to a communication protocol). In this manner, a computer-readable medium may generally correspond to (1) a tangible computer-readable storage medium that is non-transitory, or (2) a communication medium such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, codes, and / or data structures for implementing the techniques described in the present disclosure. A computer program product may include a computer-readable medium.

[0103] For example, and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage devices, magnetic disk storage devices or other magnetic storage devices, flash memory or any other medium, which can be used to store the desired program code in the form of instructions or data structures and can be accessed by a computer. Moreover, any connection is appropriately referred to as a computer-readable medium. For example, if a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL) or wireless technology (such as infrared, radio and microwave) is used to transmit instructions from a website, server or other remote source, then the coaxial cable, fiber optic cable, twisted pair, DSL or wireless technology (such as infrared, radio and microwave) is included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carriers, signals or other temporary media, but rather relate to non-temporary, tangible storage media. Disks and optical disks such as those used include compact disks (CDs), laser disks, optical disks, digital versatile disks (DVDs), floppy disks and blue-ray disks, where disks usually reproduce data magnetically, and optical disks reproduce data optically with lasers. The above combination should also be included in the scope of computer-readable media.

[0104] Instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuit systems. Therefore, the terms "processor" or "processing circuitry" as used herein may each refer to any of the foregoing structures or any other structure suitable for implementing the described techniques. In addition, in some examples, the described functionality may be provided within dedicated hardware and / or software modules. Moreover, the described techniques may be implemented entirely in one or more circuits or logic elements.

[0105] As used herein, a processing device may include a processing circuit system as described above. In some examples, the processing device may include at least one processor and at least one memory, the at least one memory having a computer code, the computer code including a set of instructions, the set of instructions, when executed by the at least one processor, causing the at least one processor to perform any of the functions described herein. In some examples, the processing device may receive the computer code including the set of instructions from at least one memory coupled to the processing device.

[0106] The techniques of the present disclosure may be implemented in a variety of devices or equipment, including wireless handsets, mobile or non-mobile computing devices, wearable or non-wearable computing devices, integrated circuits (ICs), or a set of ICs (e.g., chipsets). Various components, modules, or units are described in the present disclosure to emphasize the functional aspects of devices configured to perform the disclosed techniques, but they do not necessarily need to be implemented by different hardware units. Instead, as described above, the various units may be combined in hardware units, or provided by a collection of interoperable hardware units (including one or more processors as described above) in combination with appropriate software and / or firmware.

[0107] Aspects of the present invention are provided with reference to the following clauses:

[0108] Item 1. A method comprising: determining an entropy value of each of a plurality of data blocks by a processing circuit system of a data platform to obtain a corresponding plurality of entropy values; reorganizing the plurality of data blocks by the processing circuit system and based on the corresponding plurality of entropy values ​​to obtain a reorganized plurality of data blocks; compressing the reorganized plurality of data blocks by the processing circuit system to obtain a compressed block file; and storing the compressed block file by the processing circuit system to replace the plurality of data blocks.

[0109] Clause 2. The method as described in Clause 1 also includes: deduplicating a set of data blocks stored by the storage system through the block file manager to create a deduplicated set of data blocks; and selecting the multiple data blocks from the deduplicated set of data blocks through the processing circuit system of the data platform.

[0110] Clause 3. The method as described in Clause 1 also includes: calculating the entropy value of each data block in the set of data blocks stored by the storage system through the processing circuit system of the data platform; comparing each data block in the set of data blocks with an entropy value threshold through the block file manager; and in response to comparing each data block in the set of data blocks with the entropy value threshold, selecting the multiple data blocks from the set of data blocks based on the entropy value of each data block selected to meet the entropy value threshold through the processing circuit system of the data platform.

[0111] Clause 4. The method of clause 1, further comprising: encrypting, by a processing circuit system, the compressed block file as a single file to obtain an encrypted compressed block file.

[0112] Item 5. The method as described in Item 1 also includes: reorganizing the pointers referencing each of the multiple data blocks stored by the storage system into one of a descending order or an ascending order according to the entropy value of each of the multiple data blocks by the processing circuit system; and in response to reorganizing the pointers referencing each of the multiple data blocks into the ascending order or the descending order, updating the pointers referencing each of the multiple data blocks within the storage system into the ascending order or the descending order by the processing circuit system of the data platform.

[0113] Clause 6. A method as described in Clause 5: wherein each of the pointers referencing each of the multiple data blocks stored by the storage system is stored as a node within a linked list; and wherein the method also includes sequentially organizing the nodes within the linked list corresponding to the pointers referencing each of the multiple data blocks stored by the storage system into one of the descending order or the ascending order based on establishing a new order for the multiple data blocks.

[0114] Clause 7. The method as described in Clause 1 also includes: selecting a compression algorithm from a plurality of compression algorithms based on attributes of the plurality of data blocks; and compressing the reorganized plurality of data blocks using the selected compression algorithm to obtain the compressed block file.

[0115] Clause 8. The method of clause 1, further comprising: selecting, by the processing circuitry of the data platform, the plurality of data blocks from a data storage area for creating the block file; in response to selecting the plurality of data blocks from the data storage area for creating the block file, determining the entropy value for each of the selected plurality of data blocks by calculating the entropy value at least in part via the processing circuitry according to the following formula:

[0116] H=-1*sum(p i *log(p i ));

[0117] wherein H represents the entropy value as calculated by the processing circuitry; wherein i represents the index of each of a plurality of symbols representing each of the plurality of data blocks; wherein p i is the frequency of each of the plurality of symbols i; and wherein when the logarithm base is 2, the entropy value H is expressed in the number of bits per symbol.

[0118] Item 9. A data platform comprising: a processing circuit system; a storage system; a block file manager; a compression manager; a non-transitory computer-readable medium; and wherein when the instructions are executed by the processing circuit system, the processing circuit system is configured to: determine, by the processing circuit system, an entropy value for each of a plurality of data blocks to obtain a corresponding plurality of entropy values; reorganize, by the block file manager, the plurality of data blocks based on the corresponding plurality of entropy values ​​to obtain a reorganized plurality of data blocks; compress, by the compression manager, the reorganized plurality of data blocks to obtain a compressed block file; and store, by the storage system, the compressed block file to replace the plurality of data blocks within the storage system.

[0119] Clause 10. A data platform as described in Clause 9, wherein the instruction causes the processing circuit system to: deduplicate a set of data blocks stored by the storage system through the block file manager to create a deduplicated set of data blocks; and select the multiple data blocks from the deduplicated set of data blocks through the processing circuit system of the data platform.

[0120] Clause 11. A data platform as described in Clause 9, wherein the instructions cause the processing circuit system to: calculate the entropy value of each data block in a set of data blocks stored by the storage system through an entropy calculator of the data platform; compare each data block in the set of data blocks with an entropy value threshold through the block file manager; and in response to comparing each data block in the set of data blocks with the entropy value threshold, the instructions cause the processing circuit system to select the multiple data blocks from the set of data blocks based on the entropy value of each data block selected to satisfy the entropy value threshold.

[0121] Item 12. A data platform as described in Item 9, wherein the instruction causes the processing circuit system to: reorganize, through the block file manager of the data platform, pointers referencing each of the multiple data blocks stored by the storage system into one of a descending order or an ascending order based on the entropy value of each of the multiple data blocks; and in response to reorganizing the pointers referencing each of the multiple data blocks into the ascending order or the descending order, the instruction causes the processing circuit system to update the pointers referencing each of the multiple data blocks within the storage system into the ascending order or the descending order.

[0122] Item 13. A data platform as described in Item 12: wherein each of the pointers referencing each of the multiple data blocks stored by the storage system is stored as a node within a linked list; and wherein the instruction causes the processing circuit system to sequentially organize the nodes within the linked list corresponding to the pointers referencing each of the multiple data blocks stored by the storage system into one of the descending order or the ascending order based on establishing a new order for the multiple data blocks.

[0123] Clause 14. A data platform as described in Clause 9, wherein the instruction causes the processing circuit system to: select a compression algorithm from a plurality of compression algorithms based on the attributes of the plurality of data blocks; and use the selected compression algorithm to compress the reorganized plurality of data blocks to obtain the compressed block file.

[0124] Clause 15. The data platform of clause 9, wherein the instructions cause the processing circuit system to: select the plurality of data blocks from the storage system for use in creating the block file; in response to selecting the plurality of data blocks from the storage system, determine the entropy value for each of the selected plurality of data blocks by calculating the entropy value at least in part via the processing circuit system according to the following formula:

[0125] H=-1*sum(p i *log(p i ));

[0126] wherein H represents the entropy value as calculated by the processing circuitry; wherein i represents the index of each of a plurality of symbols representing each of the plurality of data blocks; wherein p i is the frequency of each of the plurality of symbols i; and wherein when the logarithm base is 2, the entropy value H is expressed in the number of bits per symbol.

[0127] Item 16. A computer-readable storage medium comprising instructions that, when executed, configure a processing circuit to: determine an entropy value for each of a plurality of data blocks to obtain a corresponding plurality of entropy values; reorganize the plurality of data blocks based on the corresponding plurality of entropy values ​​to obtain a reorganized plurality of data blocks; compress the reorganized plurality of data blocks to obtain a compressed block file; and store the compressed block file to replace the plurality of data blocks.

[0128] Clause 17. A computer-readable storage medium as described in Clause 16, wherein the instructions cause the processing circuit system to: deduplicate a set of data blocks stored by a storage system to create a deduplicated set of data blocks; and select the multiple data blocks from the deduplicated set of data blocks.

[0129] Clause 18. A computer-readable storage medium as described in Clause 16, wherein the instructions cause the processing circuit system to: calculate the entropy value of each data block within a set of data blocks; compare each data block within the set of data blocks with an entropy value threshold; and in response to comparing each data block within the set of data blocks with the entropy value threshold, the instructions cause the processing circuit system to select the multiple data blocks from the set of data blocks based on the entropy value of each data block being selected to satisfy the entropy value threshold.

[0130] Item 19. A computer-readable storage medium as described in Item 16, wherein the instructions cause the processing circuit system to: reorganize the pointers referencing each of the multiple data blocks stored by the storage system into one of a descending order or an ascending order based on the entropy value of each of the multiple data blocks; and in response to reorganizing the pointers referencing each of the multiple data blocks into the ascending order or the descending order, the instructions cause the processing circuit system to update the pointers referencing each of the multiple data blocks within the storage system into the ascending order or the descending order.

[0131] Clause 20. The computer-readable storage medium of clause 16, wherein the instructions cause the processing circuitry to: select the plurality of data blocks from a storage system for use in creating the block file; in response to selecting the plurality of data blocks from the storage system, the instructions cause the processing circuitry to determine the entropy value for each of the selected plurality of data blocks by calculating the entropy value at least in part via the processing circuitry according to the following formula:

[0132] H=-1*sum(p i *log(p i ));

[0133] wherein H represents the entropy value as calculated by the processing circuitry; wherein i represents the index of each of a plurality of symbols representing each of the plurality of data blocks; wherein p i is the frequency of each of the plurality of symbols i; and wherein when the logarithm base is 2, the entropy value H is expressed in the number of bits per symbol.

Claims

1. A method comprising: determining, by a processing device of the data platform, an entropy value of each of the plurality of data blocks to obtain a corresponding plurality of entropy values; reorganizing, by the processing device and based on the corresponding plurality of entropy values, the plurality of data blocks to obtain a reorganized plurality of data blocks; compressing the reorganized plurality of data blocks by the processing device to obtain a compressed block file; as well as The compressed block file is stored by the processing device to replace the plurality of data blocks.

2. The method of claim 1, further comprising: deduplicating, by the processing device, a set of data blocks stored by the storage system to create a deduplicated set of data blocks; as well as The plurality of data chunks are selected, by the processing device, from the deduplicated set of data chunks.

3. The method of claim 1, further comprising: Calculating, by the processing device, the entropy value of each data block in a set of data blocks stored by the storage system; comparing, by the processing device, each data block in the set of data blocks with an entropy value threshold; as well as In response to comparing each data block within the set of data blocks to the entropy value threshold, the plurality of data blocks are selected by the processing device from the set of data blocks based on the entropy value of each data block selected to satisfy the entropy value threshold.

4. The method of claim 1, further comprising: The compressed block file is encrypted as a single file by the processing device to obtain an encrypted compressed block file.

5. The method of claim 1, further comprising: reorganizing, by the processing device, pointers referencing each of the plurality of data blocks stored by a storage system into one of a descending order or an ascending order based on the entropy value of each of the plurality of data blocks; as well as In response to reorganizing the pointers referencing each of the plurality of data blocks into the ascending order or the descending order, updating, by the processing device, the pointers referencing each of the plurality of data blocks within the storage system into the ascending order or the descending order.

6. The method according to claim 5: wherein each of the pointers referencing each of the plurality of data blocks stored by the storage system is stored as a node within a linked list; and The method further includes sequentially organizing, by the processing device, the nodes within the linked list corresponding to the pointers referencing each of the plurality of data blocks stored by the storage system into one of the descending order or the ascending order based on establishing a new order for the plurality of data blocks.

7. The method of claim 1, further comprising: selecting, by the processing device, a compression algorithm from a plurality of compression algorithms based on the attributes of the plurality of data blocks; as well as The reorganized plurality of data blocks are compressed by the processing device using the selected compression algorithm to obtain the compressed block file.

8. The method of claim 1, further comprising: selecting, by the processing device, the plurality of data blocks from a data storage area for use in creating the block file; In response to selecting the plurality of data blocks from the data storage area for use in creating the block file, determining the entropy value for each of the selected plurality of data blocks by calculating the entropy value at least in part via the processing device according to the following formula: wherein H represents the entropy value as calculated by the processing means; wherein i represents an index of each of a plurality of symbols representing each of the plurality of data blocks; where p i is the frequency of each of the plurality of symbols i; and When the logarithm base is 2, the entropy value H is expressed in bits per symbol.

9. A data platform, comprising: Storage systems; as well as A processing device having access rights to the storage device and configured to: determining an entropy value for each of a plurality of data blocks to obtain a corresponding plurality of entropy values; reorganizing the plurality of data blocks based on the corresponding plurality of entropy values ​​to obtain a reorganized plurality of data blocks; compressing the reorganized plurality of data blocks to obtain a compressed block file; as well as The compressed block file is stored by the storage system to replace the plurality of data blocks in the storage system.

10. The data platform according to claim 9, wherein the processing device is further configured to: deduplicating a set of data chunks stored by the storage system to create a deduplicated set of data chunks; and The plurality of data chunks are selected from the deduplicated set of data chunks.

11. The data platform according to claim 9, wherein the processing device is further configured to: calculating the entropy value for each data block within a set of data blocks stored by the storage system; comparing each data block in the set of data blocks to an entropy threshold; and In response to comparing each data block within the set of data blocks to the entropy value threshold, the plurality of data blocks are selected from the set of data blocks based on the entropy value of each data block being selected to satisfy the entropy value threshold.

12. The data platform according to any one of claims 9 to 11, wherein the processing device is further configured to: reorganizing pointers referencing each of the plurality of data blocks stored by the storage system into one of a descending order or an ascending order based on the entropy value of each of the plurality of data blocks; and In response to reorganizing the pointers referencing each of the plurality of data blocks into the ascending order or the descending order, updating the pointers referencing each of the plurality of data blocks within the storage system into the ascending order or the descending order.

13. The data platform according to claim 12: wherein each of the pointers referencing each of the plurality of data blocks stored by the storage system is stored as a node within a linked list; and Wherein the processing device is further configured to sequentially organize the nodes within the linked list corresponding to the pointers referencing each of the plurality of data blocks stored by the storage system into one of the descending order or the ascending order based on establishing a new order for the plurality of data blocks.

14. The data platform according to claim 9, wherein the processing device is further configured to: selecting a compression algorithm from a plurality of compression algorithms based on attributes of the plurality of data blocks; and The reorganized plurality of data blocks are compressed using the selected compression algorithm to obtain the compressed block file.

15. A computer-readable storage medium comprising instructions which, when executed, configure one or more processors to perform the method of any one of claims 1 to 8.