File compression method and device for distributed file system

By implementing intelligent file compression methods in distributed file systems, the dual requirements of storage efficiency and cost are solved, and the storage space and cost reduction are achieved.

CN120104582APending Publication Date: 2025-06-06SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510155183.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

With the increase in the amount of data, it is difficult to meet the dual storage efficiency and cost requirements of distributed file systems.

Method used

A distributed file system file compression method is provided, including compression function switch module, request data preprocessing module, compression algorithm dynamic selection module, data compression module and data access module. Through intelligent selection of compression algorithms and dynamic adjustment of compression strategies, efficient data compression and storage can be achieved.

Benefits of technology

Significantly reduce storage space requirements, reduce storage hardware costs and maintenance costs, balance compression efficiency and decompression speed, and optimize the use of storage resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104582A_ABST
    Figure CN120104582A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of cloud computing, in particular to a distributed file system file compression method and device, and a compression function switch module opens or closes a data compression function for a specific directory or the whole file system in CephFS; the request data preprocessing module preprocesses data after receiving a read-write request of a client; the compression algorithm dynamic selection module intelligently selects the most appropriate compression algorithm according to the file characteristics and the system state; the data compression module ensures that the data is effectively compressed and stored in the distributed file system; and the data access module reads the read data before decompression in the request, and stores the data on the storage node after compressing the data in the write request. Compared with the prior art, the method has the advantages that the usability and compatibility of the system can be improved, and the technical deployment and maintenance work are simplified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cloud computing, and specifically provides a distributed file system file compression method and device. Background Art

[0002] With the rapid development of technologies such as cloud computing, big data analysis, and the Internet of Things, the amount of data generated, stored, and processed by enterprises and individuals is expanding at an unprecedented rate. This places higher demands on distributed file systems, such as CephFS, especially in terms of storage efficiency, access speed, cost control, and energy consumption.

[0003] Distributed file systems (such as CephFS) achieve high availability, scalability, and fault tolerance by distributing data across multiple nodes. However, as the amount of data increases, simply adding storage hardware can no longer meet the dual needs of efficiency and cost. Therefore, how to efficiently use storage space in a distributed file system has become an urgent problem to be solved. Summary of the invention

[0004] The present invention aims at solving the above-mentioned deficiencies of the prior art and provides a highly practical distributed file system file compression method.

[0005] A further technical task of the present invention is to provide a distributed file system file compression device that is rationally designed, safe and applicable.

[0006] The technical solution adopted by the present invention to solve its technical problem is:

[0007] A distributed file system file compression method, comprising a compression function switch module, a request data preprocessing module, a compression algorithm dynamic selection module, a data compression module and a data access module;

[0008] The compression function switch module allows users to turn on or off the data compression function for a specific directory or the entire file system in CephFS according to actual needs;

[0009] After receiving the read and write request from the client, the request data preprocessing module will preprocess the data, including segmenting the file and accurately analyzing and mapping the file offset and length in the request;

[0010] The compression algorithm dynamic selection module intelligently selects the most appropriate compression algorithm according to file characteristics and system status;

[0011] The data compression module ensures that data is effectively compressed and stored in the distributed file system;

[0012] The data access module reads data before decompression in a read request, and compresses data before saving it to a storage node in a write request.

[0013] Furthermore, in the compression function switch module, the user sets whether the files in the corresponding directory need to be compressed and saved by setting the file system to specify the directory extension attribute. When creating files in the corresponding directory subsequently, the new files or directories will inherit the extension attribute.

[0014] Furthermore, the request data preprocessing module includes:

[0015] (1) Accept client request;

[0016] (2) Compression determination;

[0017] (3) offset and length alignment;

[0018] (4)Data segmentation.

[0019] Further, in step (1), when the client issues a read or write request, the request data preprocessing module first captures and parses the request to obtain the parameters required by the client, including the file operation type, target file path, offset, and data length;

[0020] In step (2), the subsequent compression work will be performed only if the following two conditions are met, otherwise, normal data reading and writing operations are performed;

[0021] The judgment conditions include:

[0022] a. Enable extended attributes. You need to check whether the file has enabled extended attributes that support compression. Only files that are explicitly marked as compressible will be compressed.

[0023] b. File length limit: It is also necessary to determine whether the file length exceeds a preset minimum allocation unit.

[0024] Furthermore, in step (3), if the requested offset or length is not an integer multiple of the minimum allocation unit, it will be adjusted to meet this requirement;

[0025] In step (4), once the request is parsed and the offset and length are aligned, the system will then segment the requested data. Data segmentation refers to dividing the file content into a series of fixed-size data blocks, with the specified fixed size being the compression unit.

[0026] Furthermore, the compression algorithm dynamic selection module includes:

[0027] (a) pre-configured compression algorithm mapping table;

[0028] (b) file type identification;

[0029] (c) Selecting a compression algorithm;

[0030] (d) Update metadata.

[0031] Furthermore, in step (a), a detailed list is created, listing various file types and their most suitable compression algorithms;

[0032] In step (b), when a file is submitted to the system, the file type is first identified by analyzing the file extension, file header, and content to determine the file type. The file metadata is used to preliminarily determine the file type. The file type is confirmed by reading the data at the beginning of the file and checking the file header or specific pattern.

[0033] Furthermore, in step (c), according to the identified file type, a pre-configured compression algorithm mapping table is queried to find a recommended compression algorithm that matches the file type; according to the identified file type in step (b), a pre-configured compression algorithm mapping table is queried to find the most suitable compression algorithm; for file types not listed in the mapping table, a default algorithm is used or an algorithm is selected by further analyzing the characteristics;

[0034] In the step (d), after the compression algorithm is selected, the compression algorithm used is saved in the extended attributes of the file.

[0035] Furthermore, in the data compression module, it is as follows:

[0036] (1) Receiving data blocks, the data compression module receives pre-processed data blocks from the data pre-processing module, wherein the data blocks have been segmented, aligned, and are ready for compression;

[0037] (2) Initialize the compressor and initialize the compressor instance according to the selected compression algorithm;

[0038] (3) compressing data and passing the data block to the compressor for compression processing;

[0039] (4) Generate compressed metadata.

[0040] A distributed file system file compression device, comprising: at least one memory and at least one processor;

[0041] The at least one memory is used to store a machine-readable program;

[0042] The at least one processor is used to call the machine-readable program to execute a distributed file system file compression method.

[0043] Compared with the prior art, the distributed file system file compression method and device of the present invention have the following outstanding beneficial effects:

[0044] (1) The present invention significantly reduces the storage space requirement by implementing an intelligent file compression method in the CephFS distributed file system, thereby reducing storage hardware costs and maintenance fees.

[0045] (2) Intelligently select compression algorithms based on file content and dynamically adjust compression strategies based on access frequency analysis to balance compression efficiency and decompression speed and optimize storage resource usage.

[0046] (3) This application designs an API adaptation layer that is transparent to upper-layer applications and automatically handles file compression and decompression without requiring users or applications to make any changes, thereby improving ease of use and compatibility. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0048] Attached Figure 1 It is a flow chart of a distributed file system file compression method;

[0049] Attached Figure 2 It is a flow chart of a request data preprocessing module in a distributed file system file compression method;

[0050] Attached Figure 3 The invention discloses a processing flow chart of a compression algorithm dynamic selection module in a distributed file system file compression method. DETAILED DESCRIPTION

[0051] In order to enable those skilled in the art to better understand the solution of the present invention, the present invention is further described in detail below in conjunction with specific implementation methods. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0052] A best embodiment is given below:

[0053] like Figure 1-3As shown, a distributed file system file compression method in this embodiment includes a compression function switch module, a request data preprocessing module, a compression algorithm dynamic selection module, a data compression module and a data access module;

[0054] The compression switch module provides a flexible configuration option that allows users to enable or disable data compression for specific directories or the entire file system in CephFS based on actual needs. This flexibility allows administrators to dynamically adjust compression policies based on storage capacity, performance requirements, or cost considerations.

[0055] Users can set whether files in the corresponding directory need to be compressed and saved by setting the file system's specified directory extended attributes. In this way, fine-grained dynamic management of each directory in the file system can be achieved. When creating files in the corresponding directory later, the new files or directories will inherit the extended attributes.

[0056] After receiving the read and write requests from the client, the request data preprocessing module will preprocess the data, including segmenting the file and accurately analyzing and mapping the file offset and length in the request. This process prepares the necessary information for the subsequent compression operation, ensuring that the compression process is efficient and orderly, and avoiding unnecessary data duplication or repeated processing.

[0057] The main processing procedures include:

[0058] (1) Receive client request

[0059] When the client issues a read or write request, the data preprocessing module first captures and parses the request to obtain parameters such as the file operation type (read or write), target file path, offset, and data length required by the client.

[0060] (2) Compression determination

[0061] Subsequent compression work will be performed only when the following two conditions are met, otherwise, normal data reading and writing operations are performed.

[0062] The judgment conditions include:

[0063] a. Enable extended attributes: The system needs to check whether the file has enabled extended attributes that support compression. Only those files that are explicitly marked as compressible will be compressed. This gives users or administrators more control and can decide which files are suitable for compression based on the content and usage scenarios of the files.

[0064] b. File length limit: In addition to checking the extended attributes, the system also needs to determine whether the length of the file exceeds a preset minimum allocation unit. This is because for very small files, compression operations may not bring obvious storage savings, but may reduce overall performance due to the additional computational cost of compression and decompression. Setting a minimum length threshold ensures that compression operations are only applied to larger files that can really benefit from it. The minimum allocation unit is configurable.

[0065] (3) Offset and length alignment

[0066] To improve the efficiency of data processing, especially considering the subsequent data segmentation and compression operations, the system will align the requested offset and length. If the requested offset or length is not an integer multiple of the minimum allocation unit, the system will adjust it to meet this requirement. This can simplify the data processing process, avoid unnecessary data fragmentation, and improve compression efficiency.

[0067] (4) Data segmentation

[0068] Once the request is parsed and the offset and length are aligned, the system will then segment the requested data. Data segmentation refers to dividing the file content into a series of fixed-size data blocks. The specified fixed size is the compression unit. Each data block (i.e., each compression unit) can be compressed and stored independently, which not only helps improve parallel processing capabilities, but also makes data management more flexible, especially in a distributed environment, where each data block can be stored independently on different nodes, thereby achieving load balancing and redundant backup.

[0069] The compression algorithm dynamic selection module is used to intelligently select the most appropriate compression algorithm based on file characteristics and system status. The compression algorithm dynamic selection module can adapt to changing workloads and system conditions to ensure optimal compression efficiency and system performance in different scenarios. This not only saves storage space, but also maintains the system's response speed and stability.

[0070] The specific processing flow is as follows:

[0071] (1) Preconfigured compression algorithm mapping table

[0072] Create a detailed list of various file types (such as .txt, .jpg, .mp4, etc.) and their most suitable compression algorithms. This mapping table should be based on experimental data and industry standards, taking into account the compression efficiency and quality of different file types. The mapping table can be static or dynamically updated to adapt to new file formats and technological developments. Evaluate the performance of different compression algorithms on various file types, including key indicators such as compression ratio, compression speed, and decompression speed. Update the mapping table regularly based on emerging file types and algorithms to keep it current and effective.

[0073] (2) File type identification

[0074] When a file is submitted to the system, the file type is first identified. The file type can be determined by file extension, file header (magic number), content analysis, etc. The file metadata, such as the extension, is used to preliminarily determine the file type. The file type is confirmed by reading the data at the beginning of the file and checking the file header or a specific pattern (such as FF D8 FF E0 for JPEG). Deep learning models are supported, the binary stream of the input file is used, and the file type classification is output to improve the recognition accuracy.

[0075] (3) Select the compression algorithm

[0076] Based on the identified file type, query the preconfigured compression algorithm mapping table. Find the recommended compression algorithm that matches the file type. Based on the file type identified in step 2, query the preconfigured compression algorithm mapping table to find the most suitable compression algorithm. For file types not listed in the mapping table, you can use the default algorithm or further analyze its characteristics to select an algorithm.

[0077] (4) Update metadata

[0078] After selecting a compression algorithm, save the compression algorithm used in the file's extended attributes to facilitate the use of the specific compression algorithm in subsequent read and write operations.

[0079] The data compression module ensures that data is effectively compressed and stored in the distributed file system while maintaining data integrity and availability.

[0080] Here are the steps:

[0081] (1) Receive data block:

[0082] The data compression module receives pre-processed data blocks from the data pre-processing module. These data blocks have been segmented, aligned, and are ready for compression.

[0083] (2) Initialize the compressor:

[0084] Initializes the compressor instance based on the selected compression algorithm. This may involve setting compression parameters such as compression level, dictionary size, etc.

[0085] (3) Compressed data:

[0086] The data block is passed to the compressor for compression. The compressor will try to remove redundancy in the data and reduce the size of the data according to its algorithm.

[0087] (4) Generate compressed metadata:

[0088] During the compression process, some metadata may be generated, such as compression ratio, compressed data size, compression algorithm identifier, etc. This information is very important for decompression and subsequent processing.

[0089] The data access module reads data before decompression in a read request, and saves the data to the storage node after compressing the data in a write request.

[0090] Based on the above method, a distributed file system file compression device in this embodiment includes: at least one memory and at least one processor;

[0091] The at least one memory is used to store a machine-readable program;

[0092] The at least one processor is used to call the machine-readable program to execute a distributed file system file compression method.

[0093] The above-mentioned specific implementations are only specific cases of the present invention. The patent protection scope of the present invention includes but is not limited to the above-mentioned specific implementations. Any technical solutions that conform to the above-mentioned specific implementations of the present invention and any appropriate changes or substitutions made by ordinary technicians in the relevant technical field shall fall within the patent protection scope of the present invention.

[0094] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A distributed file system file compression method, characterized in that: It includes a compression function switch module, a request data preprocessing module, a compression algorithm dynamic selection module, a data compression module and a data access module; The compression function switch module allows users to turn on or off the data compression function for a specific directory or the entire file system in CephFS according to actual needs; After receiving the read and write request from the client, the request data preprocessing module will preprocess the data, including segmenting the file and accurately analyzing and mapping the file offset and length in the request; The compression algorithm dynamic selection module intelligently selects the most appropriate compression algorithm according to file characteristics and system status; The data compression module ensures that data is effectively compressed and stored in the distributed file system; The data access module reads data before decompression in a read request, and compresses data before saving it to a storage node in a write request.

2. A distributed file system file compression method according to claim 1, characterized in that: In the compression function switch module, the user sets whether the files in the corresponding directory need to be compressed and saved by setting the file system to specify the directory extension attribute. When creating files in the corresponding directory later, the new files or directories will inherit the extension attribute.

3. A distributed file system file compression method according to claim 2, characterized in that: The request data preprocessing module includes: (1) Accept client request; (2) Compression determination; (3) offset and length alignment; (4)Data segmentation.

4. A distributed file system file compression method according to claim 3, characterized in that: In step (1), when the client issues a read or write request, the request data preprocessing module first captures and parses the request to obtain the parameters required by the client, including the file operation type, target file path, offset, and data length; In step (2), the subsequent compression work will be performed only if the following two conditions are met, otherwise, normal data reading and writing operations are performed; The judgment conditions include: a. Enable extended attributes. You need to check whether the file has enabled extended attributes that support compression. Only files that are explicitly marked as compressible will be compressed. b. File length limit: It is also necessary to determine whether the file length exceeds a preset minimum allocation unit.

5. A distributed file system file compression method according to claim 3, characterized in that: In step (3), if the requested offset or length is not an integer multiple of the minimum allocation unit, it will be adjusted to meet this requirement; In step (4), once the request is parsed and the offset and length are aligned, the system will then segment the requested data. Data segmentation refers to dividing the file content into a series of fixed-size data blocks, with the specified fixed size being the compression unit.

6. A distributed file system file compression method according to claim 5, characterized in that: The compression algorithm dynamic selection module includes: (a) pre-configured compression algorithm mapping table; (b) file type identification; (c) Selecting a compression algorithm; (d) Update metadata.

7. A distributed file system file compression method according to claim 6, characterized in that: In the step (a), a detailed list is created, listing various file types and their most suitable compression algorithms; In step (b), when a file is submitted to the system, the file type is first identified by analyzing the file extension, file header, and content to determine the file type. The file metadata is used to preliminarily determine the file type. The file type is confirmed by reading the data at the beginning of the file and checking the file header or specific pattern.

8. A distributed file system file compression method according to claim 7, characterized in that: In the step (c), according to the identified file type, a pre-configured compression algorithm mapping table is queried to find a recommended compression algorithm that matches the file type; according to the identified file type in step (b), a pre-configured compression algorithm mapping table is queried to find the most suitable compression algorithm; for file types not listed in the mapping table, a default algorithm is used or an algorithm is selected by further analyzing the characteristics; In the step (d), after the compression algorithm is selected, the compression algorithm used is saved in the extended attributes of the file.

9. A distributed file system file compression method according to claim 8, characterized in that: In the data compression module, as follows: (1) Receiving a data block, the data compression module receives a pre-processed data block from the data pre-processing module, wherein the data block has been segmented, aligned, and is ready for compression; (2) Initialize the compressor and initialize the compressor instance according to the selected compression algorithm; (3) compressing data and passing the data block to the compressor for compression processing; (4) Generate compressed metadata.

10. A distributed file system file compression device, characterized in that: include: at least one memory and at least one processor; The at least one memory is used to store a machine-readable program; The at least one processor is configured to call the machine-readable program to execute the method according to any one of claims 1 to 9.