Small file storage method and device, equipment, medium and program product
By hashing the storage path of small files, the index mapping relationship between hash value and large file buckets is constructed, and large file buckets are dynamically managed, the problem of degradation of distributed file system performance in massive small file scenarios is solved, and efficient small file reading and storage is achieved.
Patent Information
- Application Number
- CN202510148484.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-02-10
AI Technical Summary
In the scenario of massive small files, the performance of distributed file systems is greatly reduced and it is difficult to meet the actual needs of commercial users. The existing technology has problems of high costs and reduced performance optimization at the hardware and software level.
By hashing the storage path of small files, the index mapping relationship between the hash value of small files and the large file bucket is constructed, and large file buckets are dynamically managed to disperse the pressure of read and write requests, and red and black trees are used to store metadata to improve retrieval efficiency.
It significantly improves the reading efficiency of small files, reduces resource consumption, improves the concurrent processing capability of distributed file systems, and ensures the stability and reliability of storage performance.
Smart Images

Figure CN120029975A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data storage, and in particular to a small file storage method, device, equipment, medium and program product. Background Art
[0002] In recent years, with the rapid development of mobile applications such as social media, short videos, and e-commerce, the number of small files generated has risen sharply. These small files are of various types, sizes, and are updated frequently. Their number can often reach tens of millions, hundreds of millions, or even billions or tens of billions, making their storage gradually become a major problem in the industry. Distributed file systems have been widely used in the field of data storage due to their advantages such as high reliability, easy scalability, and high fault tolerance. However, they mainly optimize the read and write performance of large files, and lack effective optimization methods for the read and write operations of small files. This leads to a significant increase in resource consumption and time consumption for operations such as adding, deleting, searching, and modifying small files in the scenario of massive small files, which greatly reduces the performance of distributed file systems when processing massive small files, making it difficult to meet the actual needs of commercial users.
[0003] To address this problem, most technical personnel in the industry have optimized it from both the hardware and software levels.
[0004] At the hardware level, high-performance reading and writing of small files is mainly achieved by combining high-speed solid-state drives. However, this method requires expensive hardware devices when processing massive small files, resulting in a significant increase in storage costs. In addition, hardware devices are easily affected by external environmental factors, resulting in data loss and posing a major threat to data security.
[0005] At the software level, the read and write performance of small files is optimized mainly by aggregating small files. However, this method is only applicable to the storage of a small number of small files. When the number of small files increases to a massive level, the read and write performance of small files drops sharply. The reason is that this method simply merges multiple small files into a large file, but does not deeply optimize the read and write process of small files. When reading a large number of small files, it is still necessary to traverse the aggregated large files one by one, and the target file cannot be quickly located, resulting in an exponential increase in file read and write pressure and a serious loss of read and write performance. Summary of the invention
[0006] To overcome the problems existing in the related art, the present application provides a small file storage method, device, equipment, medium and program product.
[0007] According to a first aspect of an embodiment of the present application, a small file storage method is provided, the method comprising:
[0008] Perform hash calculation on the storage path of the small file, and determine the target large file bucket where the small file is to be stored based on the obtained hash value. The target large file bucket is used to aggregate and store multiple small files. Each of the target large file buckets corresponds to a target metadata bucket and a target file name bucket. The target metadata bucket and the target file name bucket are respectively used to store the metadata and file name of the small file in the corresponding target large file bucket;
[0009] When the sum of the data volume of the small file and the data volume of the target large file bucket exceeds a preset bucket capacity threshold, a new large file bucket, a new metadata bucket and a new file name bucket are generated in the same-level directory of the target large file bucket, and part of the small files, metadata and file names stored in the target large file bucket, the target metadata bucket and the target file name bucket are correspondingly migrated to the new large file bucket, the new metadata bucket and the new file name bucket, so that the small files, the metadata of the small files and the file names of the small files are respectively stored in the target large file bucket, the target metadata bucket and the target file name bucket;
[0010] The target metadata bucket adopts a red-black tree data structure to store the metadata of the small file; storing the metadata of the small file in the target metadata bucket specifically includes: comparing the file name of the small file with the file name of the red-black tree node in the target metadata bucket in lexicographic order, and determining the storage position of the metadata of the small file in the target metadata bucket based on the comparison result;
[0011] After completing the storage of the small file, if a retrieval request for the small file is received, the target large file bucket and target metadata bucket where the small file is stored are determined based on the hash value corresponding to the small file, the metadata of the small file is retrieved from the target metadata bucket based on the file name of the small file, and the small file is extracted from the target large file bucket based on the retrieved metadata.
[0012] According to a second aspect of an embodiment of the present application, a small file storage device is provided, the device comprising:
[0013] A large file bucket determination module is used to perform hash calculation on the storage path of the small file, and determine the target large file bucket where the small file is to be stored based on the obtained hash value. The target large file bucket is used to aggregate and store multiple small files. Each of the target large file buckets corresponds to a target metadata bucket and a target file name bucket. The target metadata bucket and the target file name bucket are respectively used to store the metadata and file name of the small file in the corresponding target large file bucket;
[0014] A small file storage module is used to generate a new large file bucket, a new metadata bucket and a new file name bucket in the same-level directory of the target large file bucket when the sum of the data volume of the small file and the data volume of the target large file bucket exceeds a preset bucket capacity threshold, and migrate part of the small files, metadata and file names stored in the target large file bucket, target metadata bucket and target file name bucket to the new large file bucket, new metadata bucket and new file name bucket respectively, so as to store the small files, metadata of the small files and file names of the small files in the target large file bucket, target metadata bucket and target file name bucket respectively;
[0015] The target metadata bucket adopts a red-black tree data structure to store the metadata of the small file; storing the metadata of the small file in the target metadata bucket specifically includes: comparing the file name of the small file with the file name of the red-black tree node in the target metadata bucket in lexicographic order, and determining the storage position of the metadata of the small file in the target metadata bucket based on the comparison result;
[0016] After completing the storage of the small file, if a retrieval request for the small file is received, the target large file bucket and target metadata bucket into which the small file is aggregated are determined based on the hash value corresponding to the small file, the metadata of the small file is retrieved in the target metadata bucket based on the file name of the small file, and the small file is extracted from the target large file bucket based on the retrieved metadata.
[0017] According to a third aspect of an embodiment of the present application, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the first aspect when executing the program.
[0018] According to a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the method described in the first aspect is implemented.
[0019] According to a fifth aspect of an embodiment of the present application, a computer program product is provided, including a computer program, which implements the method described in the first aspect when executed by a processor.
[0020] The technical solution provided by the embodiments of the present application may have the following beneficial effects:
[0021] In the embodiment of the present application, by performing hash calculation on the storage path of the small file, an index mapping relationship between the hash value of the small file and the large file bucket is constructed, so that when retrieving a small file, the large file bucket where the small file is stored can be quickly located directly according to the hash value, avoiding the resource consumption caused by traversing large files one by one in the distributed file system, reducing the reading burden of small files, and significantly improving the reading efficiency of small files. At the same time, the embodiment of the present application can also monitor the storage status in the large file bucket in real time, and dynamically split the bucket and migrate the file when the amount of data in the large file bucket is close to the bucket capacity threshold, thereby effectively dispersing the read and write request pressure of a single large file bucket, adjusting the storage structure in time, avoiding single point overload, significantly improving the concurrent processing capability of the distributed file system, and ensuring the stability and reliability of storage performance.
[0022] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The accompanying drawings, which are incorporated in the specification and constitute a part of this application, illustrate embodiments consistent with the application and, together with the description, serve to explain the principles of the application.
[0024] Figure 1 It is a flowchart of a small file storage method shown in the present application according to an exemplary embodiment.
[0025] Figure 2 It is a schematic diagram of the storage of small files in a large file bucket according to an exemplary embodiment of the present application.
[0026] Figure 3 It is a schematic diagram of file migration shown in the present application according to an exemplary embodiment.
[0027] Figure 4 It is a flowchart of another small file storage method shown in the present application according to an exemplary embodiment.
[0028] Figure 5 It is a hardware structure diagram of a computer device where a small file storage device is located according to an exemplary embodiment of the present application.
[0029] Figure 6 It is a structural block diagram of a small file storage device shown in the present application according to an exemplary embodiment. DETAILED DESCRIPTION
[0030] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0031] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms of "a", "said" and "the" used in this application and the appended claims are also intended to include plural forms unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.
[0032] It should be understood that although the terms first, second, third, etc. may be used in the present application to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0033] Before discussing the present application scheme in depth, it is necessary to briefly explain some basic concepts involved in the scheme in order to better understand the content of the present application.
[0034] Small files: refers to files with a small amount of data. Taking storage performance and aggregation effects into consideration, the industry generally refers to files with a data volume of less than 8KB as small files.
[0035] Large file bucket: It can be understood as a special file or container used to aggregate and manage small files under the same storage directory. Each large file bucket can correspond to a metadata bucket and a file name bucket, which are used to store the metadata and file names of small files in the corresponding large file bucket. File metadata refers to data that can describe the file attribute characteristics and storage status, usually including the following information of the file: file name, file size, creation time, modification time, permission information, etc. In actual operation, the size of a large file bucket and its metadata bucket and file name bucket can generally be 32MB or 64MB.
[0036] Red-black tree: It is a self-balancing binary search tree with the following important characteristics: each node is either red or black; the root node is black; the leaf nodes (NIL nodes, empty nodes) are black; the two child nodes of each red node are black (there cannot be two consecutive red nodes on all paths from each leaf to the root); all paths from any node to each of its leaves contain the same number of black nodes. These characteristics ensure the self-balancing of the red-black tree, so that the height of the tree remains at a logarithmic level during insertion, deletion, and search operations, thereby improving operational efficiency.
[0037] ASCII code: is an encoding scheme based on English characters, used to represent characters in text data. ASCII code uses 7-bit binary numbers to represent 128 different characters, including uppercase and lowercase English letters, numbers 0-9, and some control characters and punctuation marks. In computer systems, ASCII code is widely used for the storage and transmission of text files and is a universal character encoding standard.
[0038] In order to solve the problems existing in the related technology, the present application provides a new small file storage method. This method performs hash calculation on the storage path of the small file, and constructs an index mapping relationship between the hash value of the small file and the large file bucket, so that when retrieving the small file, the large file bucket where the small file is stored can be quickly located directly according to the hash value, avoiding the resource consumption caused by traversing large files one by one in the distributed file system, reducing the reading burden of small files, and significantly improving the reading efficiency of small files. At the same time, the method can also monitor the storage status in the large file bucket in real time, and dynamically split the bucket and migrate the file when the amount of data in the large file bucket is close to the bucket capacity threshold, thereby effectively dispersing the read and write request pressure of a single large file bucket, adjusting the storage structure in time, avoiding single point overload, significantly improving the concurrent processing capability of the distributed file system, and ensuring the stability and reliability of storage performance.
[0039] The embodiments of the present application are described in detail below in conjunction with the accompanying drawings.
[0040] Figure 1 FIG. 1 is a flowchart of a small file storage method according to an exemplary embodiment of the present application. Figure 1 As shown, the method comprises the following steps:
[0041] S101, performing hash calculation on the storage path of the small file, and determining the target large file bucket where the small file is to be stored based on the obtained hash value, the target large file bucket being used to aggregate and store multiple small files, each target large file bucket corresponding to a target metadata bucket and a target file name bucket, the target metadata bucket and the target file name bucket being used to store metadata and file names of the small files in the corresponding target large file bucket, respectively;
[0042] S102. When the sum of the data volume of small files and the data volume of the target large file bucket exceeds a preset bucket capacity threshold, a new large file bucket, a new metadata bucket and a new file name bucket are generated in the same directory as the target large file bucket, and some small files, metadata and file names stored in the target large file bucket, the target metadata bucket and the target file name bucket are migrated to the new large file bucket, the new metadata bucket and the new file name bucket respectively, so that the small files, the metadata of the small files and the file names of the small files are stored in the target large file bucket, the target metadata bucket and the target file name bucket respectively.
[0043] In step S101, since a certain amount of resource consumption will be generated for each file stored in the distributed file system, when faced with the storage of a large number of small files, the resource consumption is huge, which leads to resource waste and performance bottlenecks. Therefore, a large file bucket can be created under the storage path of the small file, and multiple small files in the same storage directory are aggregated into the large file bucket to form a large file before storage, so as to save resource consumption, reduce the number of times the distributed file system reads and writes small files, and improve the storage performance of small files. In addition, the hash value uniquely corresponding to the small file can be obtained by performing a hash calculation on the storage path of the small file, and the large file bucket to be stored in the small file is determined by using the hash value, thereby constructing an index mapping relationship between the hash value and the large file bucket, so that when the small file is subsequently retrieved, the large file bucket where the small file is stored can be quickly located directly according to the hash value, avoiding the resource consumption generated by traversing large files one by one in the distributed file system, reducing the reading burden of small files, and significantly improving the reading efficiency of small files. Among them, the hash algorithm used for the hash calculation can be selected according to the actual needs of the user, and this application does not limit this.
[0044] Specifically, in some embodiments, in the process of determining the target large file bucket, non-character symbols in the storage path, such as " / ", "\", ":", etc., can be first removed. These symbols have no practical meaning when determining the large file bucket. Removing them can simplify the subsequent calculation process and reduce the calculation complexity. Next, the processed storage path is converted into the corresponding ASCII code value. This conversion process converts the character information into digital form, which is convenient for subsequent numerical processing and calculation. However, the data length of the converted ASCII code value is still relatively long, and may contain a lot of unnecessary information when determining the large file bucket. Therefore, in order to reduce the impact of this unnecessary information, the ASCII code value can also be processed for data dimensionality reduction. Dimensionality reduction processing can be achieved through a variety of methods, such as principal component analysis (PCA) or linear discriminant analysis (LDA), etc. These methods can extract the main features in the data and remove redundant information, thereby improving the efficiency and accuracy of data processing. After dimensionality reduction processing, the amount of data is reduced, making the hash calculation more efficient, and the obtained hash value can better reflect the essential characteristics of the small file storage path. Subsequently, the ASCII code value after dimension reduction is hashed to obtain the hash value corresponding to the small file, and the hash value is modulo processed to obtain the bucket number of the target large file bucket where the small file is to be stored. Based on the bucket number, the target large file bucket is determined. This process can map the storage path of the small file to a value within a limited range, so that this range corresponds to the number of large file buckets, and the obtained value is the bucket number of the target large file bucket where the small file is to be stored. In this way, each hash value can correspond to a unique bucket number, and then determine the specific large file bucket where the small file should be stored, realizing the function of quickly locating the large file bucket according to the hash value. In addition, the hash value has good dispersion, and can evenly distribute different small files to different buckets, avoiding the situation where some buckets are overloaded while other buckets are idle, thereby achieving load balancing and ensuring the efficient operation of the distributed file system.
[0045] In step 102, after determining the target large file bucket where the small file is to be stored, it is further possible to compare whether the sum of the data volume of the small file and the data volume of the target large file bucket exceeds the preset bucket capacity threshold to determine whether the target large file bucket has enough space to accommodate the small file currently to be stored. If the sum does not exceed the threshold, it means that the target large file bucket still has enough space, so the current small file can be stored normally in this target large file bucket, and its metadata and file name can also be stored in the corresponding target metadata bucket and target file name bucket. If the sum exceeds the threshold, it means that the target large file bucket can no longer accommodate the small file currently to be stored. At this time, a new large file bucket and its corresponding new metadata bucket and new file name bucket can be generated under the same-level directory of the target large file bucket, and then part of the data stored in the target large file bucket, target metadata bucket and target file name bucket are migrated to the corresponding new bucket to release the space of the target large file bucket, target metadata bucket and target file name bucket so that it can store the small file currently to be stored. In this way, the storage status of large file buckets can be monitored in real time, and when the amount of data in the large file bucket approaches the bucket capacity threshold, the bucket can be split and the file can be migrated dynamically, thereby effectively dispersing the read and write request pressure of a single large file bucket, adjusting the storage structure in time, avoiding single point overload, significantly improving the concurrent processing capability of the distributed file system, and ensuring the stability and reliability of storage performance.
[0046] In the process of storing small files in the target large file bucket, if the small files are not stored in the large file bucket in the order of data volume, then after deleting some small files, some scattered and discontinuous spaces, i.e., file fragments, may be generated. These file fragments will cause the subsequent storage of other files. When the file size does not match the fragment space size, additional space needs to be applied for, resulting in a waste of the original file fragment resources. Therefore, in order to avoid the waste of file fragment resources and optimize the utilization of storage space in the large file bucket, in some embodiments, the small files in the target large file bucket can be stored in sequence according to the data volume of the small files. Among them, the order of data volume can be from small to large or from large to small. The specific order can be set according to the actual needs of the user, and this application does not limit this.
[0047] Let's use a specific example to illustrate the above situation: Assume that the data volumes of small files to be stored in a large file bucket are 30, 20, 10, 5, 5, 2, 1, 1, 1, 1. If stored in sequence, the stored small files may be 30, 20, 10, 5, 5, 2, 1, 1, 1, 1 or 1, 1, 1, 1, 5, 5, 10, 20, 30; if stored in a non-sequential manner, the stored small files may be 1, 10, 1, 5, 30, 1, 2, 20, 5, 1. If the small files with data volumes of 30, 20, and 10 are subsequently deleted, a continuous space of 60 will be formed if stored in sequence; if stored in a non-sequential manner, scattered spaces of 30, 20, and 10 will be left. At this time, if there are 3 files of size 15 that need to be stored, no additional space needs to be applied for if stored in sequence, but additional space is required if stored in a non-sequential manner.
[0048] During the specific implementation process, when a small file is stored in the target large file bucket, the bucket size of the target large file bucket when the small file is not stored can be used as the starting point, and the small files can be stored in the target large file bucket in sequence according to the data volume. The stored size is the data volume of the small file, and it can also be terminated with a custom end character such as "\0" to separate the small files in the target large file bucket. Figure 2 This is a schematic diagram of a small file storage situation in a large file bucket according to an exemplary embodiment of the present application. After each small file is stored, the offset of the small file in the large file bucket can be updated to the metadata of the small file, so as to facilitate the subsequent rapid positioning and acquisition of the target small file in the target large file bucket according to the offset.
[0049] In some embodiments, the metadata of the small file may specifically include the file name of the small file, the offset of the small file in the large file bucket, the data size of the small file, the creation and modification time of the small file, the storage root node of the small file, and other information. These metadata information can fully describe the attribute characteristics and storage status of the small file, and facilitate the efficient management and retrieval of small files. Among them, the file name can uniquely identify the small file, which is convenient for using the file name to implement the indexing of the metadata of the small file; the offset and data size are helpful for quickly locating and reading the small file in the target large file bucket; the creation and modification time can provide a reference for the time attribute of the small file; and the storage root node information can indicate the node storage location of the small file in the distributed file system.
[0050] In order to achieve efficient metadata management and retrieval functions, the target metadata bucket can use a red-black tree data structure to store metadata of small files. As a self-balancing binary search tree, the self-balancing property of the red-black tree can ensure that the height of the tree is maintained at a logarithmic level, thereby maintaining a low time complexity in operations such as data insertion, deletion, and retrieval. In addition, the red-black tree occupies a small memory overhead, making it well suited for efficient retrieval and dynamic update of small file storage application scenarios.
[0051] When storing the metadata of a small file in the target metadata bucket, the file name of the small file can be compared with the file name of the red-black tree node in the target metadata bucket in lexicographic order, and the storage position of the metadata of the small file in the target metadata bucket is determined based on the comparison result. Specifically, if the lexicographic order of the file name of the small file is greater than the lexicographic order of the file name of the current red-black tree node, the right subtree of the current red-black tree node is recursively searched until there is no lexicographic order of the file name of the red-black tree node that is less than the lexicographic order of the file name of the small file, and the position where the search ends is determined as the storage position of the metadata of the small file in the target metadata bucket. Conversely, if the lexicographic order of the file name of the small file is less than the lexicographic order of the file name of the current red-black tree node, the left subtree of the current red-black tree node is recursively searched until there is no lexicographic order of the file name of the red-black tree node that is greater than the lexicographic order of the file name of the small file, and the position where the search ends is determined as the storage position of the metadata of the small file in the target metadata bucket. This storage strategy based on red-black trees can effectively ensure the orderliness of metadata of small files, optimize and improve metadata retrieval efficiency, and can quickly and accurately locate and manage the metadata of files even in the scenario of massive small files.
[0052] For the file names in the target file name bucket, in some embodiments, they can be stored in order from the earliest to the latest according to the creation time of the small files. This storage method facilitates the orderly migration of files according to the creation time when file migration is required, thereby improving the efficiency and manageability of the migration process. In the specific implementation process, each file name in the file name bucket can also be separated by a horizontal tab character such as "\t" to ensure clear distinction and easy processing of the file names.
[0053] When it is determined that some of the small files, metadata and file names stored in the target large file bucket, target metadata bucket and target file name bucket need to be migrated to the new large file bucket, new metadata bucket and new file name bucket respectively, the file names with the creation time in the second half of the target file name bucket can be migrated to the new file name bucket according to the creation time; then, according to the migrated file names, the metadata to be migrated is determined in the target metadata bucket, and the metadata to be migrated is migrated to the new metadata bucket; finally, according to the metadata to be migrated, the small files to be migrated are determined in the target large file bucket, and the small files to be migrated are migrated to the new large file bucket. Among them, the storage method of the small files that need to be migrated and their metadata and file names in the new large file bucket, new metadata bucket and new file name bucket is similar to that in the original target large file bucket, target metadata bucket and target file name bucket. For specific details, please refer to the above description, which will not be repeated here.
[0054] Through the above steps, the file migration operation is completed, and the data in the original target large file bucket, target metadata bucket and target file name bucket and the data in the newly generated new large file bucket, new metadata bucket and new file name bucket are evenly distributed, thereby releasing space in the original target large file bucket, target metadata bucket and target file name bucket, so that they can continue to store the small files to be stored and their metadata and file names, effectively avoiding overload of a single bucket. Figure 3 It is a schematic diagram of file migration shown in the present application according to an exemplary embodiment.
[0055] On this basis, since the migrated files are stored in the new large file bucket, the corresponding bucket number will change, and the bucket number is calculated by the hash value of the small file. Therefore, the hash values corresponding to these small files also need to be updated accordingly. Otherwise, when retrieving small files, the original large file bucket may be mistakenly located according to the original hash value, resulting in failure to retrieve the corresponding files. In other words, after migrating the small files to be migrated to the new large file bucket, the hash values corresponding to the small files to be migrated need to be updated according to the bucket number of the new large file bucket to ensure correct indexing and retrieval of the files.
[0056] In summary of the above embodiments, in order to enable a clearer understanding of the present application, the storage process of the small file storage method provided by the present application is illustrated below through a specific application example.
[0057] Figure 4 FIG. 1 is a flowchart of another small file storage method according to an exemplary embodiment of the present application. Figure 4 As shown, the method comprises the following steps:
[0058] S401, initialize the distributed file system. This step may specifically include: setting the mount and storage path to ensure that the file system can read and write files on the path that meets the user's needs; registering the file system information in the kernel so that the operating system can identify and manage the file system; connecting the VFS (virtual file system) interface to manage the file system through the super block, which may contain the control information of the file system, such as the size of the file system, the block size, etc.; initializing the inode root node information, which is used to describe the directory attributes of the file system and indicates the starting point of all files and directories in the file system.
[0059] S402, determine whether the currently stored file is a small file, if so, continue to execute step S403, if not, execute the normal large file storage mechanism of the distributed file system. In this step, it can be determined whether the currently stored file is a small file by comparing whether the data volume of the currently stored file meets the preset data volume range of small files. For example, if it is set that a file with a data volume less than 8k is a small file, it can be determined whether the data volume of the currently stored file is less than 8k, if so, it is determined to be a small file, if not, it is determined to be a large file.
[0060] S403: Perform a hash calculation on the storage path of the small file, and determine the target large file bucket where the small file is to be stored based on the obtained hash value.
[0061] S404, determine whether the sum of the data volume of the small file and the data volume of the target large file bucket exceeds the preset bucket capacity threshold, if so, continue to step S405, if not, directly execute step S407.
[0062] S405, generating a new large file bucket, a new metadata bucket, and a new file name bucket in the same directory as the target large file bucket;
[0063] S406. Migrate the file names with the creation time in the second half in the target file name bucket to the new file name bucket, determine the metadata to be migrated in the target metadata bucket based on the migrated file names, migrate the metadata to be migrated to the new metadata bucket, and then determine the small files to be migrated in the target large file bucket based on the metadata to be migrated, migrate the small files to be migrated to the new large file bucket, and after storing the small files to be migrated, update the offsets of the small files to be migrated in the new large file bucket to the metadata of these small files.
[0064] S407, the small file and its metadata and file name are stored in the target large file bucket, the target metadata bucket and the target file name bucket respectively, and after the small file is stored, the offset of the small file in the target large file bucket is updated to the metadata of the small file.
[0065] S408. Update the hash value corresponding to the small file to be migrated according to the bucket sequence number of the new large file bucket.
[0066] After completing the storage of small files based on the above embodiment, if a retrieval request for small files is received, the target large file bucket and target metadata bucket where the small files are stored can be determined based on the hash value corresponding to the small files, and then the metadata of the small files can be retrieved in the target metadata bucket based on the file name of the small files, and the small files can be extracted from the target large file bucket based on the retrieved metadata.
[0067] Specifically, the target large file bucket and the target metadata bucket can be determined by performing a modulo process on the hash value of the small file to obtain the bucket number of the target large file bucket where the small file is stored. Then, in the target metadata bucket, the file name of the small file is compared with the file name in the red-black tree node in lexicographic order to determine the storage location of the metadata of the small file in the target metadata bucket, thereby obtaining the metadata of the small file. Finally, based on the offset in the metadata, the storage location of the small file in the target large file bucket is determined, thereby extracting the corresponding small file from the target large file bucket. This method quickly locates the large file bucket through the hash value, achieving efficient retrieval and fast access to small files.
[0068] Corresponding to the aforementioned embodiment of the small file storage method, the present application also provides an embodiment of a small file storage device and a terminal to which it is applied.
[0069] The embodiments of the small file storage device of the present application can be applied to computer devices, such as servers or terminal devices. The device embodiments can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by its processor reading the corresponding computer program instructions in the non-volatile memory into the memory and running them. From the hardware level, if Figure 5 As shown, it is a hardware structure diagram of the computer device where the small file storage device of the embodiment of the present application is located. Figure 5 In addition to the processor 501, memory 502, network interface 503, and non-volatile memory 504 shown, the server or electronic device where the small file storage device is located in the embodiment may also include other hardware, usually according to the actual function of the computer device, which will not be described in detail.
[0070] Figure 6 1 is a structural block diagram of a small file storage device according to an exemplary embodiment of the present application. Figure 6 As shown, the device comprises:
[0071] The large file bucket determination module 601 is used to perform hash calculation on the storage path of the small file, and determine the target large file bucket where the small file is to be stored based on the obtained hash value. The target large file bucket is used to aggregate and store multiple small files. Each target large file bucket corresponds to a target metadata bucket and a target file name bucket. The target metadata bucket and the target file name bucket are respectively used to store the metadata and file name of the small file in the corresponding target large file bucket;
[0072] The small file storage module 602 is used to generate a new large file bucket, a new metadata bucket and a new file name bucket in the same directory of the target large file bucket when the sum of the data volume of the small file and the data volume of the target large file bucket exceeds a preset bucket capacity threshold, and migrate part of the small files, metadata and file names stored in the target large file bucket, the target metadata bucket and the target file name bucket to the new large file bucket, the new metadata bucket and the new file name bucket respectively, so as to store the small files, the metadata of the small files and the file names of the small files in the target large file bucket, the target metadata bucket and the target file name bucket respectively;
[0073] The target metadata bucket uses a red-black tree data structure to store the metadata of the small file; storing the metadata of the small file in the target metadata bucket specifically includes: comparing the file name of the small file with the file name of the red-black tree node in the target metadata bucket in lexicographic order, and determining the storage location of the metadata of the small file in the target metadata bucket based on the comparison result;
[0074] After completing the storage of small files, if a retrieval request for small files is received, the target large file bucket and target metadata bucket where the small files are stored are determined based on the hash value corresponding to the small files, the metadata of the small files is retrieved in the target metadata bucket based on the file name of the small files, and the small files are extracted from the target large file bucket based on the retrieved metadata.
[0075] Correspondingly, the present application also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the small file storage method described in any of the above embodiments are implemented.
[0076] Correspondingly, the present application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the small file storage method described in any of the above embodiments.
[0077] Accordingly, the present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the small file storage method recorded in any of the above embodiments.
[0078] The implementation process of the functions and effects of each module in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, which will not be repeated here.
[0079] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiment described above is only schematic, wherein the modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, that is, they may be located in one place, or they may be distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present application scheme. A person of ordinary skill in the art can understand and implement it without paying creative labor.
[0080] The above describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0081] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the inventions claimed herein. The present application is intended to cover any variations, uses or adaptations of the present application, which follow the general principles of the present application and include common knowledge or customary techniques in the art that are not claimed in the present application. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the following claims.
[0082] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
[0083] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A small file storage method, characterized in that: include: Perform hash calculation on the storage path of the small file, and determine the target large file bucket where the small file is to be stored based on the obtained hash value. The target large file bucket is used to aggregate and store multiple small files. Each of the target large file buckets corresponds to a target metadata bucket and a target file name bucket. The target metadata bucket and the target file name bucket are respectively used to store the metadata and file name of the small file in the corresponding target large file bucket; When the sum of the data volume of the small file and the data volume of the target large file bucket exceeds a preset bucket capacity threshold, a new large file bucket, a new metadata bucket and a new file name bucket are generated in the same-level directory of the target large file bucket, and part of the small files, metadata and file names stored in the target large file bucket, the target metadata bucket and the target file name bucket are correspondingly migrated to the new large file bucket, the new metadata bucket and the new file name bucket, so that the small files, the metadata of the small files and the file names of the small files are respectively stored in the target large file bucket, the target metadata bucket and the target file name bucket; The target metadata bucket adopts a red-black tree data structure to store the metadata of the small file; storing the metadata of the small file in the target metadata bucket specifically includes: comparing the file name of the small file with the file name of the red-black tree node in the target metadata bucket in lexicographic order, and determining the storage position of the metadata of the small file in the target metadata bucket based on the comparison result; After completing the storage of the small file, if a retrieval request for the small file is received, the target large file bucket and target metadata bucket where the small file is stored are determined based on the hash value corresponding to the small file, the metadata of the small file is retrieved from the target metadata bucket based on the file name of the small file, and the small file is extracted from the target large file bucket based on the retrieved metadata.
2. The method according to claim 1, characterized in that Performing a hash calculation on the storage path of the small file, and determining a target large file bucket where the small file is to be stored based on the obtained hash value, includes: Eliminate non-character symbols in the storage path of the small file, and convert the processed storage path into a corresponding ASCII code value; Performing dimension reduction processing on the ASCII code value, and performing hash calculation on the ASCII code value after the dimension reduction processing to obtain a hash value corresponding to the small file; The Hash value is modulo processed to obtain the bucket number of the target large file bucket where the small file is to be stored, and the target large file bucket is determined based on the bucket number.
3. The method according to claim 1, characterized in that The small files in the target large file bucket are stored in sequence according to the data volume of the small files.
4. The method according to claim 1, characterized in that Based on the comparison result, determining the storage location of the metadata of the small file in the target metadata bucket includes: If the lexicographic order of the file name of the small file is greater than the lexicographic order of the file name of the current red-black tree node, recursively search the right subtree of the current red-black tree node until there is no red-black tree node whose file name has a lexicographic order less than the lexicographic order of the file name of the small file, and determine the search end position as the storage position of the metadata of the small file in the target metadata bucket; If the lexicographic order of the file name of the small file is smaller than the lexicographic order of the file name of the current red-black tree node, the left subtree of the current red-black tree node is recursively searched until there is no red-black tree node whose file name has a lexicographic order greater than the lexicographic order of the file name of the small file, and the search end position is determined as the storage position of the metadata of the small file in the target metadata bucket.
5. The method according to claim 1, characterized in that The file names in the target file name bucket are stored in order from earliest to latest according to the creation time of the small files; Migrating some of the small files, metadata, and file names stored in the target large file bucket, the target metadata bucket, and the target file name bucket to the new large file bucket, the new metadata bucket, and the new file name bucket, including: According to the creation time, the file names with the creation time in the latter half of the target file name bucket are migrated to the new file name bucket; According to the file name of the migration, determine the metadata to be migrated in the target metadata bucket, and migrate the metadata to be migrated to the new metadata bucket; According to the metadata to be migrated, the small files to be migrated are determined in the target large file bucket, and the small files to be migrated are migrated to the new large file bucket.
6. The method according to claim 5, characterized in that After migrating the small files to be migrated to the new large file bucket, the method further includes: According to the bucket sequence number of the new large file bucket, the hash value corresponding to the small file to be migrated is updated.
7. A small file storage device, characterized in that: include: A large file bucket determination module is used to perform hash calculation on the storage path of the small file, and determine the target large file bucket where the small file is to be stored based on the obtained hash value. The target large file bucket is used to aggregate and store multiple small files. Each of the target large file buckets corresponds to a target metadata bucket and a target file name bucket. The target metadata bucket and the target file name bucket are respectively used to store the metadata and file name of the small file in the corresponding target large file bucket; A small file storage module is used to generate a new large file bucket, a new metadata bucket and a new file name bucket in the same-level directory of the target large file bucket when the sum of the data volume of the small file and the data volume of the target large file bucket exceeds a preset bucket capacity threshold, and migrate part of the small files, metadata and file names stored in the target large file bucket, target metadata bucket and target file name bucket to the new large file bucket, new metadata bucket and new file name bucket respectively, so as to store the small files, metadata of the small files and file names of the small files in the target large file bucket, target metadata bucket and target file name bucket respectively; The target metadata bucket adopts a red-black tree data structure to store the metadata of the small file; storing the metadata of the small file in the target metadata bucket specifically includes: comparing the file name of the small file with the file name of the red-black tree node in the target metadata bucket in lexicographic order, and determining the storage position of the metadata of the small file in the target metadata bucket based on the comparison result; After completing the storage of the small file, if a retrieval request for the small file is received, the target large file bucket and target metadata bucket where the small file is stored are determined based on the hash value corresponding to the small file, the metadata of the small file is retrieved from the target metadata bucket based on the file name of the small file, and the small file is extracted from the target large file bucket based on the retrieved metadata.
8. A computer device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method described in any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Small file storage method and device, equipment and medium
CN111176574A
Object metadata storage method and device
CN114428766A
File system management method and device, cluster, medium and product
CN119201841A
Software-Defined Network Attachable Storage System and Method
US20140082145A1
File Backup Into an Object Storage Bucket
US20250028608A1
Cited By
Data storage method and device, equipment, medium and program product
CN121166022A
Data storage method, device, equipment, medium and program product
CN121166022B