File system processing method, electronic equipment and medium
By compressing the file system data and metadata, and integrating metadata address information into the super block in memory, optimizing the disk format of the read-only file system, the problem of large-scale storage resources occupied by read-only file system is solved and reading efficiency is improved.
Patent Information
- Application Number
- CN202510549334.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-07-29
AI Technical Summary
When the existing read-only file system has a large number of files, the metadata is not compressed, resulting in the image taking up more storage resources and the number of IO and decompression times when reading metadata.
The data and metadata of files in the file system are compressed. By integrating the metadata of multiple files in memory and storing their address information into the super block, the image file of the file system is generated, and the disk format is optimized to super block + compressed metadata + compressed data.
It effectively reduces the storage space occupied by the file system, reduces the number of decompression and IO times, and reduces the time complexity of metadata reading.
Smart Images

Figure CN120386773A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of storage technology, and in particular, to a method for processing a file system, an electronic device, and a medium. Background Art
[0002] A read-only file system is a type of file system whose core feature is the immutability of data. In the disk image format of the original file system image, data blocks corresponding to metadata and data blocks corresponding to data are set for each file.
[0003] In order to reduce the storage space occupied by the file system, the file system is usually compressed. Considering that the metadata has a high access frequency and a small data volume, compression will increase the computational overhead and access latency, and the saved space is limited. Therefore, in related file system compression methods, generally only the content (i.e., data) of the file is compressed, and the metadata of each file is not compressed. However, when the number of files is large, due to the uncompressed metadata, the storage resources occupied by the image are still relatively large.
[0004] It can be seen that how to effectively reduce the storage resources occupied by the file system is a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention
[0005] The present invention provides a method for processing a file system, an electronic device, and a medium, so as to at least solve the technical problem in the related art that when the number of files in the file system is large, only the content of the file is compressed, resulting in relatively large storage resources occupied by the image.
[0006] The present invention provides a method for processing a file system, including: obtaining compressed data obtained by compressing data of files in the file system and obtaining metadata of multiple files; storing the metadata of multiple files in memory, and storing the address information of the metadata of multiple files in memory in a superblock; compressing the metadata of multiple files located in memory into one data block or multiple data blocks; generating an image file of the file system according to the superblock, the data block, and the compressed data.
[0007] The beneficial effects of the present invention are as follows. First, compared with the method of only compressing the data of files in the file system, when the number of files in the file system is large, the disk image format of the file system image is superblock + metadata of each file + data after compressing the data of each file; while in the method provided by the present invention, in addition to compressing the data of files in the file system, the metadata of files is also compressed. After generating the image file of the file system according to the superblock, data blocks and compressed data, the mirror disk format is superblock + metadata after compressing the metadata of multiple files + data after compressing the data of multiple files. In the present invention, by reorganizing the disk format of the metadata and compressing it into one or more data blocks, the size of the file system image is greatly reduced, effectively reducing the storage space occupied by the file system. Second, if both the data and metadata of files are compressed, when reading the data of files, each file needs to be decompressed separately, and the number of decompression times and IO times required is large. However, in the method provided by the present invention, the metadata of multiple files is obtained, the metadata of multiple files is stored in the memory, and the metadata of multiple files in the memory is compressed into one data block or multiple data blocks, that is, the metadata of multiple files in the file system is integrated and compressed, so that when reading the metadata of the file system, the compressed metadata of multiple files is decompressed uniformly, reducing the number of decompression times and IO times required. Third, compared with the method of using the traditional binary search to find the metadata of files, in the method provided by the present invention, the metadata of multiple files is stored in the memory, and the address information of the metadata of multiple files in the memory is stored in the superblock, so that the address of the metadata can be directly obtained when reading the metadata of the file, reducing the time complexity of metadata reading.
[0008] The present invention also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above file system processing methods when executing the computer program.
[0009] The present invention also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above file system processing methods are implemented.
[0010] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of any of the above file system processing methods are implemented. Description of the Drawings
[0011] To more clearly illustrate the embodiments of the present invention, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0012] Figure 1 Schematic diagram of an original mirror disk format provided by an embodiment of the present invention;
[0013] Figure 2 Flowchart of a method for processing a file system provided by an embodiment of the present invention;
[0014] Figure 3 Schematic diagram of an optimized disk format provided by an embodiment of the present invention;
[0015] Figure 4 Flowchart of a method for implementing metadata compression provided by an embodiment of the present invention;
[0016] Figure 5 Flowchart of a method for reading metadata provided by an embodiment of the present invention;
[0017] Figure 6 Flowchart of a method for reading a file provided by an embodiment of the present invention;
[0018] Figure 7 Schematic diagram of a device for optimizing file system storage provided by an embodiment of the present invention. Detailed implementation manners
[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.
[0020] It should be noted that in the description of the present invention, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present invention are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0021] A read-only file system is a type of file system whose core feature is the immutability of data. It allows users or programs to read the contents of files and directories, but prohibits operations such as writing to, modifying, deleting, or creating new files. It has high data security, preventing data from being accidentally modified or damaged, and is suitable for use in embedded devices, security systems, and digital media storage.
[0022] The read-only file system supports efficient data compression, which can significantly reduce the storage space occupied by the file system and save device storage resources. However, when creating the file system image, only the contents of the files are compressed, and the metadata of each file is not compressed. Thus, when there are a large number of files, the image will occupy more storage resources. Figure 1 A schematic diagram of an original image disk format provided by an embodiment of the present invention is as Figure 1 shown. The disk image format of the original file system image is superblock, metadata, data, metadata, data, metadata, data. That is, in the form of metadata + data, without compressing the metadata, which increases the storage capacity occupied by the file system image. In addition, if the file metadata and file content are directly compressed together, it will cause more input / output (IO) and data decompression when reading the file metadata information.
[0023] In order to effectively reduce the storage resources occupied by the file system, an embodiment of the present invention provides a processing method for the file system.
[0024] To enable those skilled in the art of this technology to better understand the solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Figure 2 A flowchart of a processing method for a file system provided by an embodiment of the present invention is as Figure 2 shown. The method includes:
[0025] S10: Obtain the compressed data obtained by compressing the data of the files in the file system and obtain the metadata of multiple files;
[0026] S11: Store the metadata of multiple files in memory and store the address information of the metadata of multiple files in memory in the superblock;
[0027] S12: Compress the metadata of multiple files located in memory into one data block or multiple data blocks;
[0028] S13: Generate an image file of the file system according to the superblock, data block, and compressed data.
[0029] It should be noted that the file system in the present invention specifically refers to a read-only file system. The file system contains one or more files, and each file has corresponding data (i.e., content) and metadata. The metadata of a file is descriptive information about the file, which records various attributes and characteristics of the file.
[0030] Specifically, obtaining the metadata of a file includes:
[0031] Obtaining at least the inode information, extended attribute information, and logical cluster information for maintaining the compressed data address information after compression of the file;
[0032] Taking at least the inode information, extended attribute information, and logical cluster information for maintaining the compressed data address information after compression as the metadata of the component.
[0033] In order to reduce the storage space occupied by the file system, first, the data of the files in the file system is compressed. The method of data compression is not limited and is determined according to the time situation. If only the data of the files is compressed and the metadata is not compressed, then when the number of files in the file system is large, due to the uncompressed metadata, the storage space occupied by the file system is still large. Therefore, in the present invention, the metadata of the files is further compressed.
[0034] If the metadata of each file is compressed separately, then when obtaining the metadata information of the file, it will cause more I / O and data decompression. Therefore, in the present invention, in order to further reduce the number of I / O operations and data decompression times caused when reading the metadata information of the file, the metadata of the files in the file system is sorted out, and then the metadata of the files is compressed uniformly.
[0035] Specifically, first obtain the metadata of multiple files. The metadata of the multiple files described here can refer to the metadata of all files in the file system. After obtaining the metadata of the multiple files, in order to improve the efficiency of reading the metadata of the files, apply for a piece of memory, store the metadata of the multiple files in the memory, and then store the address information of the metadata of the multiple files in the memory in the superblock. It should be noted that in the mirror storage format, the superblock is the core metadata structure of the file system, which is used to store the global information of the file system, such as the size of the file system, the size and quantity of blocks and inodes, the file system status, the mount information, and the supported features, etc. It is usually located at a specific position in the file system, and for the purpose of improving reliability, multiple redundant copies are stored. The superblock is the key to the normal operation and recovery of the file system. The operating system and file system tools obtain the configuration and status information of the file system by reading the superblock, so as to realize the management and maintenance of the file system.
[0036] For the convenience of understanding file information, the superblock also includes: a hash table for storing the file name and the unique identifier of the file;
[0037] Obtaining the hash table corresponding to the current file in the superblock includes:
[0038] Obtaining the absolute path of the current file relative to the current compressed directory, and using the absolute path as the key of the hash table;
[0039] Obtaining the unique identifier assigned to the current file, and using the unique identifier as the value of the hash table;
[0040] Mapping the key and value of the hash table into the hash table in the superblock.
[0041] In this method, by storing the hash table corresponding to the file in the superblock, that is, storing the absolute path of the file and the unique encoding of the file, the storage of file information is realized.
[0042] If the metadata of the file is compressed, in related technical solutions, after compressing the metadata of each file, the compression information (such as compression characteristics and compression algorithms) is stored in the metadata of the compressed file, resulting in the need to read the compression characteristics of the file when reading the metadata of each file, increasing the IO operations. To reduce the IO operations and the storage space occupied by the file system, the superblock also includes compression information; among them, the compression information at least includes the compression algorithm characteristics and compression algorithm types used for data compression;
[0043] Storing the compression information into the superblock includes:
[0044] Obtaining the metadata of multiple files, and removing the compression information from the metadata of multiple files;
[0045] Storing the removed compression information into the superblock.
[0046] Since the compression algorithm used by the current file system image is consistent, maintaining the compression algorithm used by the file system image in the superblock eliminates the need to read the compression characteristics of the file when reading the file, thus reducing the IO operations; and since there is no need to maintain the compression algorithm characteristics used in the metadata of each file, but instead storing the compression characteristics uniformly in the superblock, the storage space occupied by the file system is further reduced.
[0047] In order to quickly locate the address of metadata in memory, the super block also includes: an array for storing the address information of metadata in memory; the index of the array is the unique identifier of the file. By storing the address information of metadata in the form of an array and using the unique identifier of the file as the index of the array, after obtaining the unique identifier of the file, the address information of the metadata of the file in memory can be quickly determined according to the unique identifier of the file.
[0048] When storing the address information of the metadata of multiple files in memory in the super block, in order to improve the access efficiency and management accuracy of the file system for metadata, facilitate quick location and operation of metadata, and enhance the structuring and reliability of data storage. In the embodiments of the present invention, after storing the metadata of multiple files in memory and before storing the address information of the metadata of multiple files in memory in the super block, it further includes:
[0049] Obtain the page where the metadata is located, the starting address and the ending address of the metadata on the page;
[0050] Storing the address information of the metadata of multiple files in the super block includes:
[0051] Save the starting address of the metadata of multiple files in memory, the length of the metadata, and the information of the page where the metadata is located in an array.
[0052] After storing the address information of the metadata of the file in memory to the super block, compress the metadata of the file into one data block or multiple data blocks. There is no limit on the number of compressed data blocks, the compression algorithm used, etc., which is determined according to the actual situation.
[0053] In order to improve storage efficiency, optimize data management, and enhance the overall performance of the system, compressing the metadata of multiple files located in memory into one data block or multiple data blocks includes:
[0054] After compressing the metadata of multiple files located in memory, obtain the size of the compressed metadata;
[0055] Judge whether the size of the compressed metadata is an integer multiple of the data block size;
[0056] If so, generate one or more data blocks;
[0057] If not, fill in a preset number at the end of the metadata, compress the filled metadata, and return to the step of obtaining the size of the compressed metadata.
[0058] The preset number filled in is, for example, filling 0.
[0059] During the above process, after integrating the metadata of multiple files in the file system and then compressing it, in practice, there may be the same information in the metadata of different files. If the metadata is directly compressed, there will be multiple compressions of the same information, resulting in a relatively large volume of the obtained image file and relatively large storage space occupied. Therefore, in order to further reduce the storage space occupied by the image file, in implementation, compressing the metadata of multiple files in memory into one data block or multiple data blocks includes:
[0060] Obtain the same attribute information included in the metadata of multiple files in memory;
[0061] Store the same attribute information into the pre-established common attribute information;
[0062] Compress the common attribute information into one or more data blocks, and compress the remaining metadata in the metadata of multiple files except the common attribute information into one or more data blocks.
[0063] That is, when compressing, add a common inode attribute information, move the same attribute information in the inode to the common inode attribute information, and the image body can be reduced after compression. Then when reading the inode metadata, create a structure body of the inode common attribute, and the common attribute of the inode will point to this structure body.
[0064] In order to further reduce the storage space of the image file, in implementation, obtaining the compressed data obtained after data compression of the files in the file system includes:
[0065] Divide the data of the file into multiple data sub-blocks, and obtain the hash value of each data sub-block;
[0066] Obtain the data sub-blocks with the same hash value;
[0067] Retain one target data sub-block from the data sub-blocks with the same hash value;
[0068] Point the file content index in the index node of the data sub-blocks with the same hash value to the target data sub-block;
[0069] Compress the data obtained after processing the data of the files in the file system to obtain compressed data; among them, the data obtained after processing the data of the files in the file system includes the target data sub-block, the file content index, and the remaining data after removing the data sub-blocks with the same hash value from the data of the files in the file system.
[0070] There is no limit on the number of data sub-blocks divided from the data of the file, which is determined according to the actual situation of the file.
[0071] In this method, the hash value of each data sub-block is calculated, and then the hash values of each data sub-block are compared. The data sub-blocks with the same hash value are removed, and only one data sub-block is retained. The file content indexes in the inode with the same data sub-blocks are all pointed to this data block, and then the data of the file is compressed. That is, only one identical data sub-block is compressed, reducing the space occupied by the compressed data block, that is, reducing the storage space of the image file. At the same time, the compression efficiency of the file system is also improved.
[0072] It should be noted that in order to reduce the usage rate of the central processing unit and reduce the compression time, when performing data compression, an acceleration hardware can be used to unload the compression process from the central processing unit to the hardware.
[0073] In addition, in order to quickly locate the metadata after file compression, after compressing the metadata of multiple files in memory into one data block or multiple data blocks, before generating the image file of the file system according to the super block, data block, and compressed data, it further includes:
[0074] Obtain the end address of the super block;
[0075] Determine the next address adjacent to the end address of the super block according to the end address of the super block;
[0076] Start storing the data block from the next address adjacent to the end address of the super block;
[0077] Generating the image file of the file system according to the super block, data block, and compressed data includes:
[0078] Sort the super block, data block, and compressed data in descending order to generate the image file of the file system.
[0079] That is, storing the data block (i.e., the compressed metadata) after the super block enables the storage address of the compressed metadata on the disk to be quickly located according to the end address of the super block. To intuitively understand the disk format (i.e., the optimized disk format) after the file system processing method provided by the present invention, Figure 3 This is a schematic diagram of an optimized disk format provided by an embodiment of the present invention. As Figure 3 shown, the optimized disk format is: super block, metadata, data, data, data. Compared with Figure 1 the disk format, due to the integration and unified compression of the metadata of multiple files; and the compression information of multiple files is uniformly maintained in the super block, therefore, the storage space occupied by the file system is effectively reduced and the number of IO operations is reduced.
[0080] The above process is called the process of metadata compression in the processing method of the file system. To facilitate the intuitive understanding of the metadata compression process provided by the embodiments of the present invention by those skilled in the art, the process of implementing metadata compression will be described again below in combination with the accompanying drawings and specific embodiments. Figure 4 It is a flowchart of a method for implementing metadata compression provided by an embodiment of the present invention. As Figure 4 shown, the method includes:
[0081] S14: Create a read-only file system image file;
[0082] S15: When constructing the read-only file system superblock, add a hash table to the superblock;
[0083] S16: Obtain the compression algorithm characteristics and types used by the current image, and write them into the superblock information;
[0084] S17: Initialize a 4-byte array to save the metadata address information, and store it in the superblock;
[0085] S18: Traverse all files in the current environment;
[0086] S19: Obtain the file path relative to the root directory and assign a unique identifier to it;
[0087] S20: Use the file name as the key of the hash table and the unique identifier of the file as the value of the hash table, and map it to the hash table;
[0088] S21: Compress the file, and construct the inode information, extended attribute information and logical cluster information of the file as the metadata of the file;
[0089] S22: Determine whether all files have been traversed; if so, go to step S23; if not, return to step S18;
[0090] S23: Apply for a block of memory, store all the metadata in the memory, and record the start address and end address of each metadata relative to this block of memory;
[0091] S24: Put the metadata address information into the array in the superblock for storing metadata information, and use the file unique identifier as the index;
[0092] S25: Put the metadata address information into the metadata array in the superblock, and use the file unique identifier as the index;
[0093] S26: Put the start address of the metadata in the lower 12 bits of the array, and the length of the metadata in the upper 12 bits;
[0094] S27: Compress the metadata into a data block and place it after the super block;
[0095] S28: Generate a compressed image.
[0096] As Figure 4 shown, during the implementation of metadata compression:
[0097] 1) Specify a file directory and create a file system image from the specified file directory.
[0098] 2) Construct the super block information of the file system image, and add a hash table to the super block to store the file name and the unique identification ID of the file.
[0099] 3) Obtain the compression algorithm characteristics and compression algorithm type used for data compression, and store the compression information in the super block.
[0100] 4) Traverse all files in the current environment:
[0101] a. Obtain the absolute path of the current file relative to the current compression directory, assign a unique identification (Identification, ID) to the file, use the absolute path of the current file as the key of the hash table, and the unique identification ID of the file as the value of the hash table, and map it to the hash table in the super block.
[0102] b. Compress the file, and then construct the inode information of the file, the extended attribute information, and the lcluster information for maintaining the compressed data address information after compression as the metadata of the file.
[0103] 5) After traversing all files, apply for the corresponding number of pages, place the metadata of all files in memory, and record the page where each metadata is located and the start address and end address corresponding to the page.
[0104] 6) Place the address information of the metadata in the array in the super block for storing metadata information, and use the file unique identification ID as the array index. Place the start address of the metadata in the lower 12 bits of the array, the length of the metadata in the higher 12 bits, and the page where the metadata is located in the higher 8 bits.
[0105] 7) Compress the metadata into one or more data blocks and place them after the super block.
[0106] The above process realizes the compression of files in the file system (including the compression of file data and the compression of metadata). In practice, there is a need to read file metadata. Therefore, after generating the image file of the file system according to the super block, data block, and compressed data, it also includes:
[0107] Mount the image file to a preset directory and obtain the absolute path of the target file relative to the mounted directory; where the target file is the file for which metadata information is to be obtained.
[0108] Obtain the unique identifier corresponding to the target file from the superblock according to the absolute path of the target file relative to the mounted directory.
[0109] Obtain the metadata information of all files maintained by the image file.
[0110] Obtain the metadata information corresponding to the target file according to the unique identifier corresponding to the target file and according to the location where the metadata is stored and / or the location where the compressed data of the metadata is stored.
[0111] Among them, obtaining the metadata information corresponding to the target file according to the unique identifier corresponding to the target file and according to the location where the metadata is stored and / or the location where the compressed data of the metadata is stored includes three ways to obtain the metadata corresponding to the target file.
[0112] Specifically, Method 1:
[0113] Obtaining the metadata information corresponding to the target file according to the unique identifier corresponding to the target file and according to the location where the metadata is stored and / or the location where the compressed data of the metadata is stored includes:
[0114] When it is detected that the metadata is stored in memory, use the unique identifier corresponding to the target file as the index of the target array.
[0115] Obtain the address information of the metadata of the target file from the superblock according to the index of the target array.
[0116] Read the metadata information of the target file from memory according to the address information.
[0117] Method 2: Obtaining the metadata information corresponding to the target file according to the unique identifier corresponding to the target file and according to the location where the metadata is stored and / or the location where the compressed data of the metadata is stored includes:
[0118] After it is detected that the metadata is not stored in memory, if it is detected that the compressed data of the metadata is stored in memory, call the compression algorithm in the superblock to decompress the metadata.
[0119] Use the unique identifier corresponding to the target file as the index of the target array.
[0120] Obtain the address information of the metadata of the target file from the superblock according to the index of the target array.
[0121] Read the metadata information of the target file from memory according to the address information.
[0122] Method 3: Obtaining the metadata information corresponding to the target file based on the unique identifier corresponding to the target file and the location where the metadata is stored and / or the location where the compressed data of the metadata is stored includes:
[0123] After detecting that the metadata is not stored in memory, if it is detected that the compressed data of the metadata is not stored in memory, determine the address of the metadata on the disk according to the address and size of the superblock;
[0124] Read the compressed data of the metadata on the disk into memory according to the address of the metadata on the disk;
[0125] Call the compression algorithm in the superblock to decompress the metadata;
[0126] Use the unique identifier corresponding to the target file as the index of the target array;
[0127] Obtain the address information of the metadata of the target file from the superblock according to the index of the target array;
[0128] Read the metadata information of the target file from memory according to the address information.
[0129] It should be noted that after reading the metadata of the target file, in order to be able to obtain the metadata information from memory in the future, after reading the metadata information of the target file from memory according to the address information, it further includes: obtaining the metadata information; maintaining the metadata information in memory.
[0130] The above process is called the process of reading metadata in the file system's processing method. To facilitate those skilled in the art to intuitively understand the process of reading metadata provided by the embodiments of the present invention, the process of implementing the reading of metadata will be described below again in combination with the accompanying drawings and specific embodiments. Figure 5 It is a flowchart of a method for reading metadata provided by an embodiment of the present invention. As Figure 5 shown, the method includes:
[0131] S29: Obtain the characteristic information of the metadata of the file from the disk;
[0132] S30: Read the compression characteristics and compression algorithm used by the mirror from the disk;
[0133] S31: Obtain the absolute path of the file relative to the mounted directory;
[0134] S32: Obtain the superblock information of the file, and obtain the unique identifier corresponding to the file from the hash table according to the absolute path of the file;
[0135] S33: Perform the step of obtaining the metadata information of the file;
[0136] S34: Determine whether the metadata information is in the cache; if not, proceed to step S35; if so, proceed to step S39;
[0137] S35: Determine whether the compressed metadata data is in the cache; if so, proceed to step S38; if not, proceed to step S36;
[0138] S36: Calculate the address of the metadata on the disk based on the address and size of the superblock;
[0139] S37: Read the metadata into memory;
[0140] S38: Call the corresponding compression algorithm to decompress the metadata;
[0141] S39: Use the unique identifier of the file as an array index to obtain the file metadata address information from the superblock metadata array;
[0142] S40: Read the corresponding file metadata information from memory according to the address information;
[0143] S41: Maintain the metadata information in the cache when there is no metadata information in the cache.
[0144] As Figure 5 shown, in the metadata reading process:
[0145] 1) Mount the file system image to the specified directory and read the superblock information of the image.
[0146] 2) Obtain the metadata feature information of the file from the disk (such as the unique identifier of the file), obtain the absolute path of the file in the image according to the name of the file, and then, according to the file absolute path information, obtain the corresponding unique identifier ID of the file from the hash table maintained by the superblock.
[0147] 3) Obtain the metadata information of all files maintained by the image, and determine whether the metadata information is in the cache. If the metadata information is not in the cache and the compressed data of the metadata is also not in the cache, calculate the position of the metadata in the image according to the address and size of the superblock, read the metadata information from the disk, and then decompress the metadata to the page according to the compression algorithm used, and then jump to step 6).
[0148] 4) If the metadata information is not in the cache and the compressed data of the metadata is in the cache, decompress the metadata to the page according to the compression algorithm used. Then jump to step 6).
[0149] 5) If the metadata information is in the cache, directly jump to step 6).
[0150] 6) Use the unique identifier ID of the file as the data index to obtain the metadata address information of the specified file from the metadata array in the superblock. The lower 12 bits are used as the starting address of the metadata, the following 12 bits are used as the ending address of the metadata, and the higher 8 bits are used as the page offset address. Obtain the metadata information of the corresponding file from the metadata page.
[0151] The metadata is read through the above method.
[0152] In addition to the above metadata reading, there is also a process of reading files in practice. The processing methods of the file system also include:
[0153] Obtain the inode information corresponding to the file;
[0154] Determine whether the content to be read is in the cache;
[0155] If so, read the file content from the cache;
[0156] If not, read the compressed data from the disk into the memory, and obtain the compression information from the superblock; Decompress the compressed data according to the compression information.
[0157] Figure 6 It is a flowchart of a method for reading files provided by an embodiment of the present invention. As Figure 6 shown, the method includes:
[0158] S42: Read the file to be read;
[0159] S43: Obtain the corresponding inode according to the file;
[0160] S44: Determine whether the content to be read is in the cache; If so, go to step S45; If not, go to step S48;
[0161] S45: Read the compressed data from the disk;
[0162] S46: Obtain the compression algorithm and compression characteristics used for mirroring from the superblock;
[0163] S47: Decompress the compressed data according to the used compression algorithm and compression characteristics;
[0164] S48: Return the file content.
[0165] As Figure 6 shown, during the file reading process:
[0166] 1) Read the content of the specified file.
[0167] 2) Find the corresponding inode according to the file.
[0168] 3) Determine whether the read content exists in the page cache. If it exists, obtain the file content from the page cache and return the file content.
[0169] 4) If it is not in the page cache, read the compressed data from the disk into the memory, then obtain the compression algorithm and compression characteristics used by the mirror from the superblock, and decompress the compressed data according to the used compression algorithm and compression characteristics.
[0170] 5) Return the read file content.
[0171] In this method, when the content of the file is in the memory, it is directly obtained from the memory; when the content of the file is not in the memory, the compressed data on the disk is read into the memory, and then the compressed data is further decompressed in the memory to obtain the content of the file. It can be seen that the success rate of reading the content of the file is improved by the method provided in this embodiment.
[0172] In order to implement the processing method of the file system described above, in practice, a device for optimizing the storage of the file system can be set up. Figure 7 The following is a schematic diagram of a device for optimizing the storage of a file system provided by an embodiment of the present invention. As Figure 7 shown, the device for optimizing the storage of the file system includes: a metadata sorting device, a metadata compression device, a metadata reading device, a metadata caching device, and a file reading device.
[0173] Metadata sorting device: When making a mirror, after reading the metadata information of each file, remove the information maintaining the compression characteristics in each metadata and add it to the superblock. Assign a 32-bit address to each metadata, save the starting address of the file metadata to the lower 12-bit address, save the ending address of the file metadata to the higher 12-bit address, save the offset of the file metadata in the page to the higher 8-bit address, and maintain it in the superblock.
[0174] Metadata compression device: Compress the metadata into one data block or multiple data blocks. If the size of the compressed metadata is not a multiple of the data block, fill 0 at the end of the metadata to ensure that the compressed data is exactly one data block or multiple data blocks.
[0175] Metadata reading device: According to the address and size of the superblock, obtain the address of the metadata, then read the metadata from the disk, decompress the metadata according to the used compression algorithm, and finally read the metadata of the specified file according to the metadata address of the specified file.
[0176] Metadata caching device: After reading the metadata, cache the metadata. When reading the metadata next time, directly read the metadata information from the memory.
[0177] Specifically, 1) When creating a file system, construct the metadata information of each file and organize them together.
[0178] 2) Allocate a 32-bit address for each metadata. Save the starting address of the file metadata to the lower 12-bit address, save the ending address of the file metadata to the higher 12-bit address, save the offset of the file metadata in the page to the higher 8-bit address, and maintain it in the superblock.
[0179] 3) Storage of metadata data blocks. Place the compressed metadata data blocks directly after the superblock to ensure quick positioning to the metadata address.
[0180] 4) Since the compression feature used in creating the file system image remains unchanged, remove the compression feature maintained in each file metadata and maintain it in the superblock to reduce storage occupancy.
[0181] 5) When the nid of the file inode is obtained, query the one that saves the metadata address according to the nid of the file. Obtain the starting address and length of the metadata in the data block according to the higher 12 bits, lower 12 bits, and higher 8 bits of the 32-bit address, and read the metadata from the data block.
[0182] Through the combination of a metadata sorting device, a metadata compression device, a metadata reading device, and a metadata caching device, the compression and reading of metadata can be completed. By reorganizing the disk format of the metadata, compressing it into a single data block, and placing it after the superblock, and maintaining the compression algorithm used by the file system image in the superblock, the size of the file system image is reduced. At the same time, when reading a file, there is no need to read the compression feature of the file, so the IO operations are reduced; when creating the image, save the starting address and ending address of the metadata in the data block in 32-bit memory, and the address of the metadata can be directly obtained when reading the file metadata, reducing the time complexity of metadata reading.
[0183] It can be seen that the method provided by the present invention has the following advantages:
[0184] 1) The disk image format of the original file system image is in the form of metadata plus data, and the metadata is not compressed, which increases the storage capacity occupancy of the file system image. In the present invention, by reorganizing the disk format of the metadata and compressing it into a single data block, the size of the file system image is greatly reduced.
[0185] 2) When creating the file system image, place the compressed metadata data blocks directly after the superblock. Since the size and position of the superblock are fixed, when reading the metadata data blocks, the position of the data blocks on the disk can be quickly located.
[0186] 3) After reorganizing the metadata, when reading a file, if the traditional binary search is used to read the file metadata information, it will increase the time complexity and reduce the performance of the read-only file system. Here, by adding 32-bit memory for each file metadata and storing the start address and end address of the metadata in the data block in the 32-bit memory, the address of the metadata can be directly obtained when reading the file metadata, reducing the time complexity of metadata reading.
[0187] 4) Since when making the file system image, the compression algorithm characteristics used for the metadata in each file are maintained, and since the compression algorithm used by the current file system image is consistent, the compression algorithm used by the file system image is maintained in the superblock, further reducing the size of the file system image. Since there is no need to read the compression characteristics of the file when reading the file, the IO operations are reduced.
[0188] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0189] The embodiment of the present invention also provides a processing device for a file system, and the device includes:
[0190] A first acquisition module, configured to acquire the compressed data obtained by compressing the data of a file in the file system and acquire the metadata of multiple files;
[0191] A storage and placement module, configured to store the metadata of multiple files in the memory and store the address information of the metadata of multiple files in the memory in the superblock;
[0192] A compression module, configured to compress the metadata of multiple files located in the memory into one data block or multiple data blocks;
[0193] A first generation module, configured to generate an image file of the file system according to the superblock, the data block, and the compressed data.
[0194] In some embodiments, the superblock further includes: a hash table for storing the file name and the unique identifier of the file.
[0195] The processing device of the file system includes: a second acquisition module, configured to acquire the hash table corresponding to the current file in the superblock.
[0196] The second acquisition module specifically includes:
[0197] A third acquisition module, configured to acquire the absolute path of the current file relative to the current compressed directory and use the absolute path as the key of the hash table;
[0198] A fourth acquisition module, configured to acquire a unique identifier allocated to the current file, and use the unique identifier as the value of the hash table;
[0199] A mapping module, configured to map the key of the hash table and the value of the hash table into the hash table in the super block.
[0200] In some embodiments, the super block further includes compression information; wherein, the compression information at least includes the characteristics of the compression algorithm used for data compression and the type of the compression algorithm.
[0201] The processing device of the file system includes: a storage module, configured to store the compression information into the super block.
[0202] The storage module specifically includes:
[0203] An acquisition and removal module, configured to acquire the metadata of multiple files, and remove the compression information from the metadata of the multiple files;
[0204] A storage sub-module, configured to store the removed compression information into the super block.
[0205] In some embodiments, the super block further includes: an array for storing the address information of the metadata in the memory; the index of the array is the unique identifier of the file.
[0206] The processing device of the file system further includes:
[0207] A fifth acquisition module, configured to acquire the page where the metadata is located, the start address and the end address of the metadata on the page;
[0208] A first storage module, configured to store the address information of the metadata of multiple files in the memory in the super block;
[0209] The first storage module is specifically configured to store the start address of the metadata of multiple files in the memory, the length of the metadata, and the information of the page where the metadata is located in the array.
[0210] In some embodiments, the first acquisition module includes:
[0211] A sixth acquisition module, configured to acquire at least the inode information of the file, the extended attribute information, and the logical cluster information for maintaining the address information of the compressed data after compression;
[0212] A first acting module, configured to use at least the inode information, the extended attribute information, and the logical cluster information for maintaining the address information of the compressed data after compression as the metadata of the component.
[0213] In some embodiments, the compression module includes:
[0214] A seventh acquisition module, configured to acquire the size of the compressed metadata after compressing the metadata of multiple files located in the memory;
[0215] A judgment module, configured to judge whether the size of the compressed metadata is an integer multiple of the data block size; if so, trigger a second generation module; if not, trigger a padding module;
[0216] A second generation module, configured to generate one or more data blocks;
[0217] A padding module, configured to pad a preset number at the end of the metadata, compress the padded metadata, and return to trigger the seventh acquisition module.
[0218] In some embodiments, the processing device of the file system further includes:
[0219] An eighth acquisition module, configured to acquire the end address of the superblock;
[0220] A first determination module, configured to determine the next address adjacent to the end address of the superblock according to the end address of the superblock;
[0221] A second storage module, configured to store data blocks starting from the next address adjacent to the end address of the superblock;
[0222] The first generation module is specifically configured to: sort the superblock, data blocks, and compressed data in descending order to generate an image file of the file system.
[0223] In some embodiments, the processing device of the file system further includes:
[0224] A mounting and acquisition module, configured to mount the image file to a preset directory and acquire the absolute path of the target file relative to the mounted directory; wherein, the target file is the file for which metadata information is to be acquired;
[0225] A ninth acquisition module, configured to acquire the unique identifier corresponding to the target file from the superblock according to the absolute path of the target file relative to the mounted directory;
[0226] A tenth acquisition module, configured to acquire the metadata information of all files maintained by the image file;
[0227] An eleventh acquisition module, configured to acquire the metadata information corresponding to the target file according to the unique identifier corresponding to the target file and according to the metadata storage location and / or the location where the compressed data of the metadata is stored.
[0228] In some embodiments, the eleventh acquisition module specifically includes:
[0229] A second acting module, configured to use the unique identifier corresponding to the target file as the index of the target array when it is detected that the metadata is stored in the memory;
[0230] A first sub-acquisition module, configured to acquire the address information of the metadata of the target file from the superblock according to the index of the target array;
[0231] A first reading module, configured to read the metadata information of the target file from the memory according to the address information.
[0232] In some embodiments, the eleventh acquisition module specifically includes:
[0233] A first calling module, configured to call the compression algorithm in the superblock to decompress the metadata if it is detected that the compressed data of the metadata is stored in the memory after it is detected that the metadata is not stored in the memory;
[0234] A third acting module, configured to use the unique identifier corresponding to the target file as the index of the target array;
[0235] A second sub-acquisition module, configured to acquire the address information of the metadata of the target file from the superblock according to the index of the target array;
[0236] A second reading module, configured to read the metadata information of the target file from the memory according to the address information.
[0237] In some embodiments, the eleventh acquisition module specifically includes:
[0238] A second determination module, configured to determine the address of the metadata on the disk according to the address of the superblock and the size of the superblock if it is detected that the compressed data of the metadata is not stored in the memory after it is detected that the metadata is not stored in the memory;
[0239] A third reading module, configured to read the compressed data of the metadata on the disk into the memory according to the address of the metadata on the disk;
[0240] A second calling module, configured to call the compression algorithm in the superblock to decompress the metadata;
[0241] A fourth acting module, configured to use the unique identifier corresponding to the target file as the index of the target array;
[0242] A third sub-acquisition module, configured to acquire the address information of the metadata of the target file from the superblock according to the index of the target array;
[0243] A fourth reading module, configured to read the metadata information of the target file from the memory according to the address information.
[0244] In some embodiments, the processing device of the file system further includes:
[0245] A twelfth acquisition module, configured to acquire metadata information;
[0246] A maintenance module, configured to maintain the metadata information in memory.
[0247] For the description of the features in the corresponding embodiments of the processing device of the file system, reference can be made to the relevant descriptions in the corresponding embodiments of the file system processing method, which will not be elaborated here one by one.
[0248] An embodiment of the present invention further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above-described embodiments of the file system processing method.
[0249] An embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any one of the above-described embodiments of the file system processing method when running.
[0250] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical disks, etc., various media that can store computer programs.
[0251] An embodiment of the present invention further provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any one of the above-described embodiments of the file system processing method.
[0252] An embodiment of the present invention further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any one of the above-described embodiments of the file system processing method.
[0253] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.
[0254] The above has introduced in detail a method for processing a file system, an electronic device, and a medium provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.
Claims
1. A method for processing a file system, characterized in that including: obtaining compressed data obtained by compressing data of a file in a file system and obtaining metadata of multiple files; storing the metadata of the multiple files in memory and storing address information of the metadata of the multiple files in memory in a superblock; compressing the metadata of the multiple files located in memory into one data block or multiple data blocks; generating an image file of the file system according to the superblock, the data block, and the compressed data.
2. The processing method of the file system according to claim 1, characterized in that, The superblock further includes: a hash table for storing a file name and a unique identifier of a file; Obtaining a hash table corresponding to a current file in the superblock includes: obtaining an absolute path of the current file relative to a current compressed directory and using the absolute path as a key of the hash table; obtaining a unique identifier assigned to the current file and using the unique identifier as a value of the hash table; mapping the key of the hash table and the value of the hash table into the hash table in the superblock.
3. The processing method of the file system according to claim 2, characterized in that, The superblock further includes compression information; wherein, the compression information at least includes compression algorithm characteristics and compression algorithm types used for data compression; Storing the compression information into the superblock includes: obtaining the metadata of the multiple files and removing the compression information from the metadata of the multiple files; storing the removed compression information into the superblock.
4. The method for processing a file system according to claim 3, wherein, The superblock further includes: an array for storing address information of metadata in memory; an index of the array is a unique identifier of a file; After storing the metadata of the multiple files in memory and before storing address information of the metadata of the multiple files in memory in the superblock, it further includes: obtaining a page where the metadata is located, a start address and an end address of the metadata on the page; Storing the address information of the metadata of the multiple files in memory in the superblock includes: storing a start address of the metadata of the multiple files in memory, a length of the metadata, and information of a page where the metadata is located in the array.
5. The processing method of the file system according to any one of claims 1 to 4, characterized in that, Obtaining metadata of a file includes: at least obtaining inode information of the file, extended attribute information, and logical cluster information for maintaining address information of compressed data after compression; using at least inode information, extended attribute information, and logical cluster information for maintaining address information of compressed data after compression as metadata of a component.
6. The processing method of the file system according to any one of claims 1 to 4, characterized in that, Compressing the metadata of the multiple files located in memory into one data block or multiple data blocks includes: after compressing the metadata of the multiple files located in memory, obtaining a size of the compressed metadata; judging whether the size of the compressed metadata is an integer multiple of a data block size; if so, generating one or more data blocks; if not, padding a preset number at the end of the metadata, compressing the padded metadata, and returning to the step of obtaining the size of the compressed metadata.
7. The processing method of the file system according to claim 6, wherein Compressing the metadata of the multiple files located in memory into one data block or multiple data blocks includes: obtaining identical attribute information included in the metadata of the multiple files located in memory; storing the identical attribute information into pre-established common attribute information; Compress the common attribute information into one or more data blocks, and compress the remaining metadata in the metadata of the multiple files except the common attribute information into one or more data blocks.
8. The processing method of the file system according to claim 1, characterized in that The obtained compressed data after compressing the data of the file in the file system includes: Divide the data of the file into multiple data sub-blocks, and obtain the hash value of each data sub-block; Obtain the data sub-blocks with the same hash value; Reserve a target data sub-block from the data sub-blocks with the same hash value; Point the file content index in the inode of the data sub-blocks with the same hash value to the target data sub-block; Compress the data obtained after processing the data of the file in the file system to obtain compressed data; wherein, the data obtained after processing the data of the file in the file system includes the target data sub-block, the file content index, and the remaining data after removing the data sub-blocks with the same hash value from the data of the file in the file system.
9. The processing method of the file system according to claim 4, characterized in that, After compressing the metadata of the multiple files located in the memory into one data block or multiple data blocks, before generating the image file of the file system according to the super block, the data block, and the compressed data, it further includes: Obtain the end address of the super block; Determine the next address adjacent to the end address of the super block according to the end address of the super block; Start storing the data block from the next address adjacent to the end address of the super block; The generating the image file of the file system according to the super block, the data block, and the compressed data includes: Sort the super block, the data block, and the compressed data in descending order to generate the image file of the file system.
10. The method for processing a file system according to claim 9, characterized in that, After generating the image file of the file system according to the super block, the data block, and the compressed data, it further includes: Mount the image file to a preset directory, and obtain the absolute path of the target file relative to the mounted directory; wherein, the target file is the file for which metadata information is to be obtained; Obtain the unique identifier corresponding to the target file from the super block according to the absolute path of the target file relative to the mounted directory; Obtain the metadata information of all files maintained by the image file; Obtain the metadata information corresponding to the target file according to the unique identifier corresponding to the target file and according to the storage location of the metadata and / or the storage location of the compressed data of the metadata.
11. The method for processing a file system according to claim 10, wherein The obtaining the metadata information corresponding to the target file according to the unique identifier corresponding to the target file and according to the storage location of the metadata and / or the storage location of the compressed data of the metadata includes: When it is detected that the metadata is stored in the memory, use the unique identifier corresponding to the target file as the index of the target array; Obtain the address information of the metadata of the target file from the super block according to the index of the target array; Read the metadata information of the target file from the memory according to the address information.
12. The processing method of the file system according to claim 10, wherein, The obtaining the metadata information corresponding to the target file according to the unique identifier corresponding to the target file and according to the storage location of the metadata and / or the storage location of the compressed data of the metadata includes: After detecting that the metadata is not stored in memory, if it is detected that the compressed data of the metadata is stored in memory, the compression algorithm in the superblock is called to decompress the metadata; Use the unique identifier corresponding to the target file as the index of the target array; Obtain the address information of the metadata of the target file from the superblock according to the index of the target array; Read the metadata information of the target file from memory according to the address information.
13. The method for processing a file system according to claim 10, wherein The obtaining the metadata information corresponding to the target file according to the unique identifier corresponding to the target file and according to the storage location of the metadata and / or the storage location of the compressed data of the metadata includes: After detecting that the metadata is not stored in memory, if it is detected that the compressed data of the metadata is not stored in memory, determine the address of the metadata on the disk according to the address and size of the superblock; Read the compressed data of the metadata on the disk into memory according to the address of the metadata on the disk; Call the compression algorithm in the superblock to decompress the metadata; Use the unique identifier corresponding to the target file as the index of the target array; Obtain the address information of the metadata of the target file from the superblock according to the index of the target array; Read the metadata information of the target file from memory according to the address information.
14. An electronic device, characterized in that, Includes: A memory for storing a computer program; A processor for implementing the steps of the processing method of the file system according to any one of claims 1 to 13 when executing the computer program.
15. A computer-readable storage medium, wherein, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the processing method of the file system according to any one of claims 1 to 13 when executed by a processor.
Citation Information
Cited By
Reconstruction method of read-only file system mirror image, electronic equipment and storage medium
CN120872914A
A method for reconstructing a read-only file system image, an electronic device, and a storage medium
CN120872914B