File storage method and device, storage medium and electronic equipment

By dynamically allocating stripe sizes and using merging techniques, the problems of wasted storage resources and metadata service bottlenecks in traditional file storage systems are solved, thereby improving storage efficiency and performance.

CN120631839BActive Publication Date: 2025-11-25INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511117063.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-11-25
Estimated Expiration
2045-08-11

AI Technical Summary

Technical Problem

The fixed stripe size allocation strategy in traditional distributed file storage systems leads to low storage efficiency, especially when processing small files, resulting in wasted storage space and frequent access to metadata services.

Method used

By obtaining the business type and file type of the target file, a suitable stripe size is dynamically allocated using a machine learning-based stripe prediction model. Combined with stripe merging technology, the allocation and management of storage resources are optimized.

Benefits of technology

It significantly improves storage space utilization, reduces the frequency of metadata service access, and enhances the overall efficiency and performance of file storage, especially when dealing with small files and rapidly growing files.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120631839B_ABST
    Figure CN120631839B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a file storage method and device, a storage medium and an electronic device, and relate to the field of computers, and the method comprises the following steps: in response to a file storage request of a target file, obtaining a service type and a file type of the target file; performing strip size prediction on the target file according to the service type and the file type, to obtain a reference strip size of the target file; and allocating a target strip with the reference strip size to the target file, wherein the target strip is used for storing the target file.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computers, and more specifically, to a file storage method and apparatus, a storage medium, and an electronic device. Background Technology

[0002] In traditional distributed file storage systems, the fixed stripe size allocation strategy leads to inefficient storage, especially when dealing with small files. Specifically, for files smaller than the stripe size, a fixed stripe allocation strategy will allocate a full stripe, resulting in a large amount of storage space being filled with unused zero data. This not only wastes valuable storage resources but also increases network overhead and latency due to frequent metadata service (MDS) accesses, further contributing to low file storage efficiency.

[0003] Therefore, the related technologies suffer from the technical problem of low file storage efficiency. Summary of the Invention

[0004] This application provides a file storage method and apparatus, storage medium and electronic device to at least solve the technical problem of low efficiency in file storage in related technologies.

[0005] According to one embodiment of this application, a file storage method is provided, comprising: in response to a file storage request for a target file, obtaining the business type and file type of the target file; predicting the stripe size of the target file based on the business type and file type to obtain a reference stripe size of the target file; and allocating a target stripe of the reference stripe size to the target file, wherein the target stripe is used to store the target file.

[0006] According to another embodiment of this application, a file storage device is provided, comprising: an acquisition unit, configured to acquire the business type and file type of the target file in response to a file storage request for the target file; a prediction unit, configured to predict the strip size of the target file based on the business type and file type to obtain a reference strip size of the target file; and an allocation unit, configured to allocate a target strip of the reference strip size to the target file, wherein the target strip is used to store the target file.

[0007] According to yet another embodiment of this application, a computer-readable storage medium is also provided, in which a computer program is stored, wherein the computer program is configured to perform the steps in any of the above method embodiments when it is run.

[0008] According to yet another embodiment of this application, an electronic device is also provided, including a memory and a processor, wherein a computer program is stored in the memory and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.

[0009] The embodiments provided in this application, in response to a file storage request for a target file, obtain the business type and file type of the target file, then predict the stripe size of the target file based on the business type and file type to obtain a reference stripe size for the target file, and allocate a target stripe of the reference stripe size to the target file. Through a dynamic stripe allocation strategy, stripes are dynamically allocated according to the actual business type and file type of the file, effectively avoiding storage waste when small files occupy large stripes. This effectively solves the problems of storage resource waste, metadata service bottlenecks, and rigid stripe management caused by the fixed stripe size allocation mechanism in traditional distributed file storage systems, significantly improving the utilization rate of storage space, thereby achieving the technical effect of improving file storage efficiency and solving the technical problem of low file storage efficiency in related technologies. Attached Figure Description

[0010] Figure 1 This is a hardware structure block diagram of a file storage method according to an embodiment of this application;

[0011] Figure 2 This is a flowchart of a file storage method according to an embodiment of this application;

[0012] Figure 3 This is a schematic diagram of a file-granular dynamic stripe allocation method according to an embodiment of this application;

[0013] Figure 4 This is a schematic diagram of a stripe allocation system architecture according to an embodiment of this application;

[0014] Figure 5 This is a structural block diagram of a file storage device according to an embodiment of this application. Detailed Implementation

[0015] The embodiments of this application will be described in detail below with reference to the accompanying drawings and examples.

[0016] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0017] The methods and embodiments provided in this application can be executed on a computer terminal or similar computing device. Taking running on a computer terminal as an example, Figure 1 This is a hardware structure block diagram of a computer terminal for a file storage method according to an embodiment of this application. For example... Figure 1 As shown, a computer terminal may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 (which may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.) and a memory 104 for storing data are also shown. The computer terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the computer terminal described above. For example, the computer terminal may also include components that are more complex than those described above. Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0018] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the mapping relationship determination method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer programs stored in the memory 104, thus implementing the above-described method. The memory 104 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to a computer terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0019] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by a communication provider for the computer terminal. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module used for wireless communication with the Internet.

[0020] As an alternative solution, a file storage method, such as Figure 2 As shown, the specific steps include:

[0021] S202, in response to the file storage request for the target file, obtain the business type and file type of the target file;

[0022] S204, Based on the business type and file type, predict the strip size of the target file to obtain the reference strip size of the target file;

[0023] S206, allocate a target stripe of reference stripe size to the target file, wherein the target stripe is used to store the target file.

[0024] Optionally, in this embodiment, Business Type refers to the application scenario to which the file or data belongs, such as log recording, database operation, multimedia file storage, etc. Business Type helps predict the future growth trend and access patterns of files. File Type indicates the format or category of the file, such as text file, image file, video file, etc. File Type usually affects its size and storage requirements.

[0025] Optionally, in this embodiment, stripe size prediction is a process of predicting the optimal stripe size based on the business type and file type of the file. This can be achieved, but is not limited to, by analyzing historical I / O patterns and application characteristics. The reference stripe size is a stripe size suitable for a specific file, derived from the prediction results, aiming to reduce storage waste and improve read / write efficiency.

[0026] Optionally, in this embodiment, when a client initiates a file storage request, the system first captures this event and extracts the target file's business type and file type information. For example, if a client is uploading an image file as a user avatar, the storage system needs to identify that this is an image file belonging to the user interface management business type. This identification process can be based on various methods such as filename, extension, metadata, or content analysis.

[0027] Next, using the acquired business type and file type information, the built-in stripe size prediction model is used to predict the most suitable stripe size for the target file. The prediction model can be based on machine learning algorithms, such as Long Short-Term Memory (LSTM) networks, which learn from historical file I / O patterns to predict future growth trends for different file types. For example, for an image file expected to be no larger than 1MB, the model might predict an optimal stripe size of 1MB, while for a log file, due to its continuously growing nature, the model might suggest an initial stripe size of 4MB.

[0028] For example, suppose we are handling a video streaming service where video clips are uploaded as multiple small files. Based on past experience, we know that these video clips are typically pieced together into a complete video file after uploading. Therefore, for the first uploaded video clip, even if its initial size is only a few hundred KB, the prediction model might give a large reference stripe size, such as 4 MB, because the file is expected to grow to this size in the future. This way, during subsequent data writing, there is no need to frequently reallocate stripes, avoiding data migration and performance degradation.

[0029] The system allocates appropriate target stripes based on the predicted reference stripe size. A target stripe refers to the physical storage unit selected by the system based on the prediction results for actually storing the target file. For example, if the prediction model suggests a 1MB stripe, the system will search for a free 1MB stripe in the storage pool and allocate it to the target file. This allocation process considers the stripe distribution strategy to ensure data security and redundancy.

[0030] Continuing with the example above, if a 4MB stripe is pre-allocated to a video segment, then as subsequent data is written, as long as the total data volume does not exceed 4MB, no additional stripe allocation or data migration will be triggered. In this way, even if the video segment eventually grows to several megabytes or even tens of megabytes, high storage efficiency and read / write performance can be maintained.

[0031] Understandably, the stripe size prediction model intelligently selects the appropriate stripe size for different files. This design breaks the limitations of fixed stripe sizes in traditional distributed file systems. Especially when dealing with small files and rapidly growing files, it can significantly improve storage space utilization, reduce the frequency of metadata service access, and avoid stripe fragmentation, thereby improving the overall performance of the storage system. For example, in scenarios involving a large number of small files, this embodiment can avoid the "strip filling waste" problem, i.e., the waste of space by small files occupying large stripes; while when dealing with continuously growing files, it can optimize the storage layout in real time through dynamic stripe adjustment, reducing unnecessary data migration.

[0032] By intelligently predicting and dynamically adjusting stripe size, and through efficient metadata management, key pain points in traditional file storage systems, such as storage waste, performance bottlenecks, and rigid stripe management, are effectively addressed. This significantly improves the system's storage space utilization and read / write performance, making it particularly suitable for large-scale file storage applications in big data and cloud computing environments.

[0033] The embodiments provided in this application, in response to a file storage request for a target file, obtain the business type and file type of the target file, then predict the stripe size of the target file based on the business type and file type to obtain a reference stripe size for the target file, and allocate a target stripe of the reference stripe size to the target file. Through a dynamic stripe allocation strategy, stripes are dynamically allocated according to the actual size of the file, effectively avoiding storage waste when small files occupy large stripes. This effectively solves the problems of storage resource waste, metadata service bottlenecks, and rigid stripe management caused by the fixed stripe size allocation mechanism in traditional distributed file storage systems, significantly improving the utilization rate of storage space, thereby achieving the technical effect of improving file storage efficiency.

[0034] As an optional approach, based on the business type and file type, the stripe size of the target file is predicted, and the reference stripe size of the target file includes:

[0035] The business type and file type are input into the strip prediction model to predict the strip size, and the reference strip size output by the strip prediction model is obtained. The strip prediction model is a neural network model obtained by deep learning based on historical sample data of historical files. The historical sample data includes the historical business type, the historical file type, and the historical strip size corresponding to the historical file.

[0036] Optionally, in this embodiment, the stripe prediction model is a deep learning-based neural network model that learns from sample data of historical files to predict the stripe size to be assigned to a new file. The historical sample data includes the business type, file type, and corresponding historical stripe size of the historical files. The model learns the relationship between file size and stripe size using this data.

[0037] Optionally, in this embodiment, historical sample data refers to relevant information about files already stored in the system, including the file's business type, file type, and the actual stripe size occupied by the file. This data is used to train the stripe prediction model, enabling it to accurately predict the stripe size of new files.

[0038] Optionally, in this embodiment, when a file storage request is received from a client, the system first needs to analyze the attributes of the target file, including its business type and file type. This is the basis for predicting the stripe size, because different business types and file types often correspond to different file sizes and growth patterns.

[0039] The obtained business type and file type information are input into the stripe prediction model. This model is a neural network model that is trained using deep learning based on a large amount of historical sample data. It can predict the reference stripe size that matches the target file. The historical sample data includes the business type, file type, and actual stripe sizes used in historical files. This allows the model to accurately predict the stripe size requirements of new files by learning the patterns and rules in the historical data.

[0040] The stripe prediction model uses deep learning algorithms to analyze the input business type and file type information, and outputs a reference stripe size as the basis for stripe allocation when the target file is initially stored. Determining the reference stripe size helps avoid stripe filling waste and rigid stripe management, thereby improving storage efficiency and performance.

[0041] It should be noted that in existing technologies, the stripe size during file storage is often fixed. This "one-size-fits-all" strategy is inefficient for files of varying sizes, especially small files, where a fixed large stripe size leads to a large amount of unused zero-padding, wasting storage space. Furthermore, the fixed stripe strategy lacks flexibility as files grow; each time a file exceeds the current stripe size, it needs to be reallocated, increasing data migration overhead and network latency. To address these issues, this embodiment proposes a deep learning-based stripe size prediction technique.

[0042] The stripe prediction model accurately predicts the stripe size requirements of new files by learning from historical files' business types, file types, and actual stripe sizes. The model is trained using historical sample data, which includes various attributes of files already stored in the system, such as business type (logs, databases, images, etc.), file type (txt, jpg, mp4, etc.), and the actual stripe size occupied by the files. Through deep learning, the model can capture the implicit relationship between file characteristics and stripe size, providing a reasonable reference stripe size for new files.

[0043] In practice, when a client sends a file storage request to the system, the system first analyzes the business type and file type of the file, and then inputs this information into the stripe prediction model. The model outputs a predicted stripe size as a reference size for the target file.

[0044] The embodiments provided in this application enable more efficient allocation of storage resources, reduce zero-fill waste, and avoid performance bottlenecks caused by frequent stripe reallocation, thereby significantly improving storage space utilization and overall system read / write performance. This method leverages model prediction capabilities to achieve fine-grained management of storage resources and optimizes stripe allocation strategies, making it an effective way to improve storage efficiency and reduce overhead in distributed storage systems.

[0045] As an optional approach, the business type and file type are input into the stripe prediction model to predict the stripe size, resulting in a reference stripe size output by the model, including:

[0046] If the business type and file type are the first occurrences of the data types, obtain the initial size of the target file;

[0047] Obtain the target size interval to which the initial size belongs among N size intervals, where the size of the i-th size interval among the N size intervals is equal to the size of the first i-1 size intervals, N is a positive integer greater than 2, and i is a positive integer greater than 1 and less than N;

[0048] The upper limit of the target size range is determined as the reference strip size.

[0049] Optionally, in this embodiment, when encountering a combination of business type and file type appearing for the first time in the system, the method will not immediately rely on the prediction results of the model prediction engine. This is because the prediction model may lack sufficient training data to make accurate predictions when facing a completely new type for the first time. In this case, the system will revert to a more basic strategy—directly obtaining the initial size of the target file.

[0050] Based on N pre-defined size intervals, the interval to which the initial size of the target file belongs is located. The design of the N size intervals follows a progressive principle, meaning that the size of the i-th size interval is a specific multiple of the total size of the previous i-1 size intervals. This method ensures that the selection of strip sizes is more precise and reasonable, avoiding over-allocation or under-allocation.

[0051] Once the initial size range of the target file is found, the upper limit of that range is used as the reference strip size. This strategy takes into account both the current size of the file and reserves some room for growth, avoiding frequent stripe reallocation.

[0052] The embodiments provided in this application allow for the determination of the size of a newly appearing file based on its initial size, rather than by directly processing it using a stripe prediction model. This avoids the inaccuracy of the stripe prediction model in predicting new types of files and improves the accuracy of stripe size selection.

[0053] As an optional approach, after determining the upper limit of the target size range as the reference strip size, the method further includes:

[0054] Add business type, file type, and reference strip size to the historical sample data;

[0055] The strip prediction model is updated using supplemented historical sample data.

[0056] Optionally, in this embodiment, each determination of the strip size is recorded, forming part of the historical sample data. This data will be used to further train the strip prediction model, enabling it to make more accurate predictions when encountering the same or similar situations in the future.

[0057] Through the embodiments provided in this application, the stripe allocation strategy becomes more intelligent and efficient through dynamic adjustment and self-learning mechanisms, which can adapt to the storage needs of files of various sizes and types, thereby significantly improving the overall storage efficiency and performance of the distributed file system.

[0058] As an optional approach, after assigning a target strip of reference strip size to the target file, the method further includes:

[0059] In the case where the target stripe includes at least two stripes of the reference stripe size, the overall stripe utilization of the at least two stripes is obtained, wherein the overall stripe utilization is used to indicate the ratio of the file size stored in the at least two stripes to the stripe size of the at least two stripes;

[0060] If the overall strip utilization rate is greater than a preset threshold, strips that meet the expected conditions from at least two strips are merged.

[0061] Optionally, in this embodiment, the overall stripe utilization rate refers to the ratio between the total amount of data stored in multiple stripes (at least two) allocated to the same file and their total physical size. The overall stripe utilization rate is a key indicator for measuring stripe storage efficiency and space utilization, used to determine whether stripe merging is necessary.

[0062] Optionally, in this embodiment, the preset threshold is one of the set stripe merging trigger conditions, typically 80%. When the overall utilization of multiple stripes occupied by a file exceeds this threshold, the system will consider performing stripe merging to reduce storage system fragmentation and improve storage efficiency and read / write performance.

[0063] Optionally, in this embodiment, when the dynamically allocated reference strip size target strip includes at least two stripes, the system calculates the overall utilization of these stripes. This calculation process involves collecting file size information stored within each stripe and comparing it with the physical size of the stripe to arrive at an overall utilization value.

[0064] Once the overall stripe utilization is calculated, the system compares it to a preset threshold. If the overall stripe utilization exceeds the preset threshold (e.g., 80%), it indicates that the currently allocated stripes may not be sufficient to efficiently store the file, and there is a need to merge the stripes.

[0065] If the overall stripe utilization exceeds a preset threshold, the system will further check the expected merging conditions, such as whether the stripes are nearly fully utilized or whether they were generated during the writing process of the same file. Only when all expected conditions are met will the stripe merging operation be triggered, merging multiple stripes into a larger stripe to reduce storage system fragmentation and improve storage space utilization and read / write performance.

[0066] The embodiments provided in this application calculate the overall stripe utilization rate, which is the ratio of the total amount of stored data to the physical size of the stripe. If the overall stripe utilization rate reaches a high preset threshold (e.g., 80%), it indicates that the currently allocated stripes may no longer be sufficient to efficiently store files, and there is room for optimization by merging multiple stripes to reduce fragmentation and improve space utilization. Next, it checks whether the expected merging conditions are met, such as whether the number of stripes is sufficient and whether the stripes are compatible for merging. Once all conditions are met, the system performs a stripe merging operation, merging multiple small stripes into a larger stripe, automatically expanding by powers of 2 to reduce stripe fragmentation, improve the read / write performance and storage space utilization of the storage system for small to medium-sized files, and avoid additional overhead and write latency fluctuations caused by data migration. This mechanism, combining dynamic stripe merging with lightweight metadata indexing, can significantly improve the overall performance and resource utilization of the storage system without increasing the burden of metadata services.

[0067] As an alternative approach, the method further includes the following steps before merging strips that meet the expected conditions from at least two strips:

[0068] Obtain the stripe utilization of each stripe in at least two stripes, where the stripe utilization is used to indicate the ratio of the file size stored in each stripe to the stripe size of each stripe;

[0069] The stripes whose utilization rate is greater than a preset threshold but less than the full utilization threshold are identified as stripes that meet the expected conditions.

[0070] Optionally, in this embodiment, stripe utilization is a local concept. For a single stripe, it refers to the ratio between the effective data size of the file stored within that stripe and the physical allocated size of that stripe. A higher stripe utilization indicates that the storage space within the stripe is utilized more fully, while a lower utilization means that there are more unused blank areas.

[0071] Optionally, in this embodiment, the preset threshold is a threshold preset in the stripe merging determination criteria, typically set to 80%, used to determine whether a certain stripe needs to be merged with other stripes. When the stripe utilization exceeds this threshold, it means that the data within the stripe is close to saturation, and the merging operation can effectively reduce storage system fragmentation and improve read / write performance and space utilization.

[0072] It should be noted that the strips to be merged are strips whose utilization rate exceeds the preset threshold and is close to saturation. That is, the strip utilization rate is greater than the preset threshold but less than the full threshold, where the full threshold is usually 100%.

[0073] Optionally, in this embodiment, before preparing to merge at least two stripes, the system first calculates the utilization rate of each stripe participating in the merging candidate, that is, the ratio of the amount of file data stored inside to its physical size.

[0074] Based on the calculation results of individual stripe utilization, the system will filter out those stripes with utilization rates greater than a preset threshold (e.g., 80%) that are close to saturation. These stripes are considered to meet the expected merging conditions. The preset threshold is selected based on the system's consideration of storage efficiency and performance balance, aiming to identify those stripes with high storage efficiency that can significantly reduce fragmentation after merging.

[0075] By selecting stripes that meet preset criteria, the system identifies candidate stripes for merging. These stripes will then be merged in the next step to improve storage resource utilization efficiency and reduce read / write latency.

[0076] The embodiments provided in this application filter which stripes are most suitable candidates for merging operations by calculating the utilization rate of individual stripes. First, the system obtains the utilization rate of each stripe in at least two stripes and determines whether it reaches a preset high threshold (e.g., 80%). This individual checking approach ensures that each merging operation improves the overall efficiency of the storage system, rather than blindly merging which could lead to efficiency degradation due to unreasonable merging. In this way, only stripes with high utilization rates that significantly reduce storage fragmentation, improve space utilization, and enhance read / write performance after merging are marked as "meeting the expected conditions." This step is a further refinement based on the overall stripe utilization judgment, ensuring that the stripe merging operation considers not only the overall utilization rate but also the state of individual stripes, making the merging operation more accurate and efficient. Ultimately, through this fine-grained management and adaptive adjustment mechanism, the distributed file storage system can more intelligently manage storage resources, cope with the complex storage needs of files of different sizes and types, and significantly improve system performance and user experience.

[0077] As an optional approach, obtain the overall strip utilization of at least two stripes, including:

[0078] Get a first accumulated value of the file size stored in each of at least two stripes, and get a second accumulated value of the stripe size of each stripe;

[0079] The result of dividing the first accumulated value by the second accumulated value is determined as the overall strip utilization rate.

[0080] Optionally, in this embodiment, the first accumulated value refers to the sum of the sizes of all stored files within a set of candidate stripes (at least two) during strip merging analysis. It reflects the total space occupied by the file data and is an important numerical basis for calculating the overall strip utilization rate.

[0081] Optionally, in this embodiment, the second accumulated value is relative to the first accumulated value, and the second accumulated value refers to the sum of the physical dimensions of this group of candidate stripes. The sum of physical dimensions represents the total amount of space allocated to the system for storage of the stripes.

[0082] Optionally, in this embodiment, the first step in calculating the overall stripe utilization is to determine the total size of all stored files within the candidate stripe, i.e., the first accumulated value. This involves traversing each stripe, collecting the amount of file data stored in them, and summing them up.

[0083] Next, a second cumulative value is calculated, which is the sum of the physical dimensions of all candidate stripes. This step provides basic information about the overall storage space allocated to stripes.

[0084] After obtaining the first and second accumulated values, the system divides the first accumulated value (total file size) by the second accumulated value (total physical size), and the result is regarded as the overall stripe utilization rate. This ratio intuitively reflects the average storage density of file data in the candidate stripes.

[0085] The embodiments provided in this application calculate the total size of all stored files within these stripes (first accumulated value) and the total physical size of the stripes (second accumulated value). These two values ​​form the basis of stripe utilization analysis, used to assess the space usage of stripes when storing file data. Subsequently, by dividing the first accumulated value by the second accumulated value, the system obtains a ratio reflecting the overall storage efficiency of the stripes, namely, the overall stripe utilization rate. This calculation process ensures that the evaluation of stripe utilization considers the overall performance of all stripes, not just the condition of a single stripe. A higher overall stripe utilization rate means better average fill rate of file data in these stripes, and higher system space utilization. This indicator is one of the key criteria for deciding whether to initiate stripe merging operations. Only when the overall stripe utilization rate exceeds a preset threshold (e.g., 80%) does the system consider stripe merging beneficial, reducing storage fragmentation and improving read / write performance. In this way, the present invention provides a method for accurately evaluating stripe storage efficiency, laying a solid foundation for intelligent stripe management and helping distributed file storage systems achieve more efficient and flexible storage resource management.

[0086] As an optional approach, obtaining a first accumulated value of the file size stored within each of at least two stripes, and obtaining a second accumulated value of the stripe size for each stripe, including:

[0087] The file storage status corresponding to each stripe is obtained from the stripe index table. The file storage status is used to indicate the effective data length of the file stored in each stripe and the stripe size of each stripe. The effective data length is used to determine the first accumulated value. The stripe index table is an index table that is maintained for each stripe and records the file storage status in real time.

[0088] Optionally, in this embodiment, the stripe index table is a lightweight data structure maintained by the metadata service (MDS) to record key information such as the effective data length and stripe size of each stripe. The stripe index table is updated in real time, ensuring the system's accurate understanding of the stripe status and serving as a crucial foundation for implementing dynamic stripe allocation and merging strategies.

[0089] Optionally, in this embodiment, the file storage status indicator records real-time information about the stripe in the stripe index table, mainly including the effective data length of the stripe (i.e., the actual size of the stored file data) and the stripe size (the size of the physical allocation). Through this information, the system can understand the storage efficiency and usage of the stripe, providing a basis for subsequent stripe utilization calculations.

[0090] Optionally, in this embodiment, before calculating stripe utilization, the necessary information, including the effective data length and stripe size of each stripe, is extracted from the stripe index table. This data forms the basis for evaluating stripe utilization and reflects the actual storage situation of the stripes.

[0091] The effective data lengths of each stripe obtained from the index table are summed to form the first accumulated value; similarly, the stripe sizes of each stripe are also summed to form the second accumulated value. This summation process ensures the comprehensiveness and accuracy of the merge operation decision.

[0092] In the embodiments provided in this application, the stripe index table plays a core role, recording and updating the storage status of stripes in real time, including the actual length of file data and the physical allocation size of the stripes. By accessing the stripe index table, the system can quickly obtain the file storage status information of all candidate stripes (at least two), and then calculate the first accumulated value (the sum of the effective data length of all stored files) and the second accumulated value (the sum of the physical sizes of all candidate stripes). This process reflects refined management of storage resources, ensuring that the system can make a judgment based on accurate and comprehensive data before deciding whether to merge stripes. The existence of the stripe index table not only simplifies the data collection process and improves computational efficiency, but also provides strong support for subsequent calculation of overall stripe utilization and merging decisions.

[0093] As an alternative, the aforementioned file storage method can be applied to a scenario where distributed file storage uses dynamic stripe allocation based on file granularity. In this scenario, distributed file systems face efficiency bottlenecks when processing small files. Existing technologies use a fixed 4MB stripe size, resulting in "stripe filling waste" for files smaller than the stripe size. For example, a 1KB file might occupy the entire 4MB stripe, leading to low storage space utilization. Furthermore, the fixed stripe strategy cannot dynamically adjust with file growth; each file write requires accessing the metadata service (MDS) to obtain the stripe position, resulting in frequent network interactions and increased latency.

[0094] To address the aforementioned issues, this embodiment provides a dynamic stripe allocation method based on file granularity. For example... Figure 3 As shown, the file size is used to allocate stripes in a hierarchical manner. When the client receives a file write request, an initial stripe is generated based on an AI prediction engine (such as the stripe prediction model mentioned above). For unknown business and file types, the AI ​​engine will not be able to match initially (it will be able to identify this type after subsequent training). The first write of this type uses the initial file size (S) for judgment.

[0095] If S≤128KB, allocate a 128KB stripe;

[0096] If 128KB < S ≤ 256KB, allocate a 256KB stripe;

[0097] If 256KB < S ≤ 512KB, allocate a 512KB stripe;

[0098] If 512KB < S ≤ 1024KB, allocate a 1MB stripe;

[0099] If 1MB < S ≤ 2MB, allocate a 2MB stripe;

[0100] If S > 2MB, allocate a 4MB stripe.

[0101] Optionally, the AI prediction engine can be, but is not limited to, the above stripe prediction model, which will not be elaborated here.

[0102] As an optional example, the pseudocode of the AI prediction engine is as follows:

[0103] # Pseudocode example: File growth prediction based on LSTM

[0104] def predict_stripe_size(file_type,history_io_pattern):

[0105] if file_type == "LOG" and history_io_pattern.append_ratio > 80%:

[0106] return 1024KB # Pre-allocate a large stripe [[ID=​​​​​​​​​​​​​Triggering stripe merging conditions:

[0112] Automatic strip merging trigger conditions: The parameters are explained below:

[0113] n: The number of original stripes corresponding to this file.

[0114] T: Merging threshold (Recommended value: 0.8, i.e., 80%).

[0115] L i : The effective data length of the i-th stripe (in bytes).

[0116] S i Physical size of the i-th stripe (in bytes).

[0117] Target strip size after merging: The parameters are explained below:

[0118] S new : Size of the target strip after merging (in KB).

[0119] S max The maximum stripe size supported by the system is currently set to 4096KB (4MB) by default.

[0120] S sum : The physical size occupied by the strips to be merged.

[0121] [x]: Round x up.

[0122] Optionally, in this embodiment, an index table is maintained for each stripe in the MDS to record the effective data offset, stripe status, and the stripe size it occupies.

[0123] The index table structure is shown in Table 1, as follows:

[0124] Table 1 Index Table

[0125]

[0126] Explanation of variables in each column:

[0127] Stripe ID: A unique identifier for each stripe, used to locate and manage stripes in a distributed system. Value range: 64-bit unsigned integer (0x0000000000000000 to 0xFFFFFFFFFFFFFFFF).

[0128] Data Start Offset: This refers to the starting position of valid data within a stripe, offset from the stripe's start address, measured in bytes. It is used to skip metadata or padding data in the stripe header and directly locate the starting position of valid data. Value range: 0 ≤ Offset ≤ (Stripe size – 1).

[0129] Effective Data Length: Refers to the actual size of valid data stored within the stripe, measured in bytes. It is used to accurately record the amount of data and avoid reading invalid padding data. Value range: 0 ≤ length ≤ stripe size.

[0130] Stripe Size: The amount of space allocated to a stripe in physical storage, dynamically selected by the system based on the file size. Value range: hierarchical stripe sizes (128KB, 256KB, 512KB, 1MB, 2MB, 4MB).

[0131] Stripe Status: Describes the current usage status of the stripe and provides operational recommendations for stripe management decisions. Value range: Underutilized: Utilization < 80%. Fully Utilized: Utilization ≥ 100%. Mergeable: 80% ≤ Utilization < 100% and close to the upper limit.

[0132] Optionally, in this embodiment, a schematic diagram of a stripe allocation system architecture is shown below. Figure 4 As shown, the client module uses the fuse client, which includes a stripe pre-allocation engine and supports caching stripe metadata for the 100 most recently accessed files. It also implements small input / output (IO) aggregation, merging multiple small requests for the same file into a single stripe write. The Metadata Service (MDS) module adds a dynamic stripe allocator, dynamically selecting the stripe level based on file size; it maintains tiered stripe pools (128KB pool, 256KB pool, etc.) and allocates OSD storage locations using the CRUSH algorithm. The OSD module supports sparse stripe storage and automatically skips zero-filled regions.

[0133] The embodiments provided in this application can improve space utilization: Initial stripes (128KB~4MB) are dynamically selected based on file size, breaking the limitations of traditional fixed stripes and significantly reducing storage waste for small files (e.g., a 1KB file is reduced from occupying 4MB to 128KB). Stripe fragmentation can be avoided: Merging is triggered by a mathematical formula (utilization threshold T≥0.8 and n≥2), expanding in powers of 2 to avoid fragmentation, which is superior to traditional periodic defragmentation (e.g., Ceph's Staggered scrub), offering higher real-time performance and reduced computational overhead. Metadata access performance can be improved: The index table records effective data offsets and stripe status, combined with tiered stripe pools (128KB / 256KB, etc.), improving metadata query efficiency. Storage read / write performance can be improved: Dynamic stripe matching reduces data migration and storage performance degradation.

[0134] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.

[0135] This embodiment also provides a file storage device for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0136] Figure 5 This is a structural block diagram of a file storage device according to an embodiment of this application, such as... Figure 5 As shown, the device includes:

[0137] The acquisition unit 502 is used to acquire the business type and file type of the target file in response to the file storage request of the target file;

[0138] Prediction unit 504 is used to predict the strip size of the target file based on the business type and file type, and obtain the reference strip size of the target file;

[0139] Allocation unit 506 is used to allocate a target strip of reference strip size to the target file, wherein the target strip is used to store the target file.

[0140] As an optional solution, prediction unit 504 includes:

[0141] The prediction module is used to input the business type and file type into the strip prediction model to predict the strip size and obtain the reference strip size output by the strip prediction model. The strip prediction model is a neural network model obtained by deep learning based on historical sample data of historical files. The historical sample data includes the historical business type of the historical file, the historical file type of the historical file, and the historical strip size corresponding to the historical file.

[0142] As an optional solution, the prediction module includes:

[0143] The first acquisition submodule is used to obtain the initial size of the target file when the business type and file type are the first occurrences of the data types;

[0144] The second acquisition submodule is used to acquire the target size interval to which the initial size belongs in N size intervals, where the size of the i-th size interval in the N size intervals is equal to the size of the first i-1 size intervals, N is a positive integer greater than 2, and i is a positive integer greater than 1 and less than N;

[0145] The first determination submodule is used to determine the upper limit value of the interval corresponding to the target size interval as the reference strip size.

[0146] As an optional solution, the device also includes:

[0147] The supplementary module is used to supplement the historical sample data with the business type, file type, and reference strip size after the upper limit value of the target size range is determined as the reference strip size;

[0148] The update module is used to update the strip prediction model using supplemented historical sample data after the upper limit value of the target size range is determined as the reference strip size.

[0149] As an optional solution, the device also includes:

[0150] The first acquisition module is used to acquire the overall stripe utilization of at least two stripes after allocating a target stripe of reference stripe size to a target file, provided that the target stripe includes at least two stripes of reference stripe size, wherein the overall stripe utilization is used to indicate the ratio of the file size stored in the at least two stripes to the stripe size of the at least two stripes.

[0151] The merging module is used to merge at least two stripes that meet the expected conditions after allocating target stripes of reference strip size to the target file, provided that the overall stripe utilization rate is greater than a preset threshold.

[0152] As an optional solution, the device also includes:

[0153] The second acquisition module is used to acquire the stripe utilization rate of each stripe in at least two stripes before merging stripes that meet the expected conditions in at least two stripes, wherein the stripe utilization rate is used to indicate the ratio of the file size stored in each stripe to the stripe size of each stripe.

[0154] The determination module is used to identify the stripes with a strip utilization rate greater than a preset threshold as stripes that meet the expected conditions before merging the stripes that meet the expected conditions from at least two stripes.

[0155] As an optional solution, the first acquisition module includes:

[0156] The third acquisition submodule is used to acquire the first accumulated value of the file size stored in each of the at least two stripes, and to acquire the second accumulated value of the stripe size of each stripe;

[0157] The second determination submodule is used to determine the overall strip utilization rate by dividing the first accumulated value by the second accumulated value.

[0158] As an optional solution, the third acquisition submodule includes:

[0159] The acquisition sub-unit is used to retrieve the file storage status corresponding to each stripe from the stripe index table. The file storage status is used to indicate the effective data length of the file stored in each stripe and the stripe size of each stripe. The effective data length is used to determine the first accumulated value. The stripe index table is an index table maintained for each stripe and records the file storage status in real time.

[0160] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.

[0161] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.

[0162] It should be noted that the above modules can be implemented by software or hardware. For the latter, they can be implemented in the following ways, but are not limited to: all the above modules are located in the same processor; or, the above modules are located in different processors in any combination.

[0163] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above method embodiments when run.

[0164] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0165] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.

[0166] In one exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.

[0167] Embodiments of this application also provide a computer program product, including a non-volatile computer-readable storage medium storing the computer program product, wherein the computer program, when executed by a processor, implements the steps of the methods in various embodiments of this application.

[0168] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.

[0169] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.

[0170] The above are merely preferred embodiments of this application and are not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.

Claims

1. A file storage method, characterized in that, include: In response to a file storage request for a target file, obtain the business type and file type of the target file; Based on the business type and the file type, the strip size of the target file is predicted to obtain the reference strip size of the target file; Assign a target strip of the reference strip size to the target file, wherein the target strip is used to store the target file; When the target stripe includes at least two stripes of the reference stripe size, the file storage status corresponding to each stripe is obtained from the stripe index table, wherein the file storage status is used to indicate the effective data length of the file stored in each stripe and the stripe size of each stripe, and the stripe index table is an index table maintained for each stripe and recording the file storage status in real time; Based on the file storage status, obtain a first cumulative value of the file size stored in each of the at least two stripes, and obtain a second cumulative value of the stripe size of each stripe; The result of dividing the first accumulated value by the second accumulated value is determined as the overall stripe utilization rate of the at least two stripes, wherein the overall stripe utilization rate is used to indicate the ratio of the file size stored in the at least two stripes to the stripe size of the at least two stripes; When the overall stripe utilization rate is greater than a preset threshold, a stripe that meets the expected conditions is determined from the at least two stripes. The stripe that meets the expected conditions is a stripe whose stripe utilization rate is greater than the preset threshold and less than the full threshold. The stripe utilization rate is used to indicate the ratio of the file size stored in the stripe to the stripe size. Obtain the strip size and value of the strip that meets the expected conditions; The base-2 logarithm corresponding to the strip size and value is determined as the first parameter. The first parameter is rounded up to obtain the second parameter. The second parameter raised to the power of 2 is determined as the strip size to be merged. The strips that meet the expected conditions are merged, wherein the size of the merged strip is the smaller of the following parameters: the preset maximum strip size and the size of the strip to be merged.

2. The method according to claim 1, characterized in that, The step of predicting the stripe size of the target file based on the business type and the file type to obtain the reference stripe size of the target file includes: The business type and the file type are input into the strip prediction model to predict the strip size, and the reference strip size output by the strip prediction model is obtained. The strip prediction model is a neural network model obtained by deep learning based on historical sample data of historical files. The historical sample data includes the historical business type of the historical file, the historical file type of the historical file, and the historical strip size corresponding to the historical file.

3. The method according to claim 2, characterized in that, The step of inputting the business type and the file type into the stripe prediction model, predicting the stripe size, and obtaining the reference stripe size output by the stripe prediction model includes: If the business type and the file type are data types that appear for the first time, obtain the initial size of the target file; Obtain the target size interval to which the initial size belongs among N size intervals, wherein the size of the i-th size interval among the N size intervals is equal to the size of the first i-1 size intervals, N is a positive integer greater than 2, and i is a positive integer greater than 1 and less than N; The upper limit of the target size range is determined as the reference strip size.

4. The method according to claim 3, characterized in that, After determining the upper limit of the interval corresponding to the target size interval as the reference strip size, the method further includes: The business type, the file type, and the reference strip size are added to the historical sample data; The strip prediction model is updated using the supplemented historical sample data.

5. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the method according to any one of claims 1 to 4.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • File storage using variable stripe sizes

    CN106233264A