Progressive data striping method for task-level burst cache file system

By adopting a progressive data striping method in the burst cache file system, differentiated striping processing of dynamic adjustment of files is solved, and the problem of low data processing efficiency in traditional file systems in high-performance computing jobs is achieved, achieving higher throughput and response speed.

CN120122893AActive Publication Date: 2025-06-10NAT UNIV OF DEFENSE TECH
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510601072.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-06-10
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

Traditional burst cache file systems are difficult to efficiently process when processing large-scale data and burst access loads in high-performance computing jobs, resulting in unbalanced cache resource utilization and may cause cache overflow or performance bottlenecks.

Method used

The progressive data striping method is adopted to differentiate the striping of different parts of the file, and the file block size and storage node number are dynamically adjusted according to the file size and data access mode, thereby realizing progressive data striping.

Benefits of technology

By dynamically adjusting the strip size and number of nodes, unnecessary I/O requests and delays are reduced, data storage performance and access efficiency are improved, and system throughput and response speed are improved. It is suitable for application scenarios with high concurrency and large data volume.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120122893A_ABST
    Figure CN120122893A_ABST
Patent Text Reader

Abstract

The invention discloses a progressive data striping method for a task-level burst cache file system, which comprises the following steps of: configuring a striping component to carry out progressive data striping on a file, initializing the file, and carrying out data striping on the file, the file system operates the striped file; the progressive data striping means that all parts with different file sizes are segmented by using different file block sizes, and different node numbers are used for configuring read-write operation; and from the head of the file, the size of each part is monotonically increased, the size of the used file block is monotonically increased, and the number of nodes used for read-write operation is also monotonically increased. The method aims at improving the data storage and access performance of the task-level burst cache file system and improving the data access efficiency and reliability in high-performance computing operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data management of a burst cache file system, and particularly relates to a progressive data striping method for a task-level burst cache file system. Background Art

[0002] In high-performance computing (HPC) jobs, a large amount of data needs to be processed and analyzed, and storage also faces more complex patterns, including a large number of metadata operations, small I / O requests, or random file I / Os, and high-speed data access is required to meet the needs of real-time computing and analysis. To improve data storage efficiency, burst cache technology has been widely used. A burst cache file system is usually a separate file system that can store temporary data for applications. It aggregates the high-speed local storage of computing nodes or uses a dedicated solid-state drive storage cluster to provide a peak higher than the backend bandwidth, reduce I / O operation latency, and increase data throughput, thereby accelerating job applications.

[0003] However, with the increase in the data volume of HPC jobs and the complexity of computing tasks, traditional burst cache file systems face the problem of being unable to efficiently process large-scale data and burst access loads. Existing cache management strategies usually adopt fixed cache strategies or static data striping techniques, which may lead to unbalanced utilization of cache resources and even cache overflow or performance bottlenecks under high loads. Therefore, how to improve the throughput of data access in high-performance computing jobs has become an urgent problem to be solved. To improve data throughput and optimize cache resource utilization, existing methods mainly include the following three: One is to use a fixed striping strategy. Its basic idea is to allocate data to multiple storage media according to a fixed strip size, so as to achieve parallel access. This method is suitable for systems with uniform loads and simple data access patterns and can improve throughput through parallel reading and writing. However, it is difficult to cope with dynamic loads and complex data access patterns. For scenarios with burst loads or frequently changing access patterns, fixed striping may cause some storage devices to be overloaded while other devices are idle, thus wasting storage resources. At the same time, the choice of strip size directly affects performance, but for different computing tasks and data access scenarios, choosing a fixed strip size is often not optimal. The second is to use cache prefetching and cache management strategies. At present, some high-performance computing systems have adopted cache prefetching technology to reduce I / O latency by loading the data that may be accessed in advance into the cache. Combined with intelligent cache management, it can predict and schedule data access according to the data access pattern, thereby improving the cache hit rate and data throughput. Due to relying on the accurate prediction of the data access pattern, if the access pattern is unstable or difficult to predict, the prefetching strategy may backfire, resulting in cache pollution and resource waste, and it also has high requirements for the utilization of cache resources. If the cache resources are limited, the effect of the prefetching strategy may not be as expected. The third is to use a multi-level cache architecture by setting different levels of caches (for example, level-1 cache, level-2 cache, etc.) to more efficiently process data access requests at different levels. This method optimizes the data access path through a reasonable hierarchical structure and improves data throughput. However, such system architectures are complex and require managing multiple cache levels and data migration strategies. In addition, there are high requirements for tuning the cache levels and cache strategies. The cache sizes and usage strategies at different levels must be reasonably configured to achieve the optimal performance. In summary, existing methods can optimize the data access efficiency in high-performance computing jobs, but their respective limitations may also affect the overall efficiency and reliability of the system in certain specific scenarios. Summary of the Invention

[0004] The technical problem to be solved by the present invention: Aiming at the above problems of the prior art, a progressive data striping method for a task-level burst cache file system is provided. The present invention aims to improve the data storage and access performance of the task-level burst cache file system and enhance the data access efficiency and reliability in high-performance computing jobs.

[0005] To solve the above technical problems, the technical solution adopted by the present invention is as follows: A progressive data striping method for a task-level burst cache file system, comprising the following steps: configuring a striping component to perform progressive data striping on a file, initializing the file, and the file system operating on the striped file; the progressive data striping means that each part with a different file size is separately sliced using different file block sizes and configured to use different numbers of nodes for read and write operations, and starting from the head of the file, the size of each part monotonically increases, the file block size used monotonically increases, and the number of nodes used for read and write operations also monotonically increases.

[0006] Optionally, when performing progressive data striping, it includes slicing the part of the file from 0 to 4MB using a file block size of 512KB, and using 1 node to process the read and write operations; slicing the part of the file from 4MB to 8MB using a file block size of 1MB, and using 3 nodes to process the read and write operations; slicing the part of the file from 8MB to 16MB using a file block size of 2MB, and using 8 nodes to process the read and write operations; slicing the part of the file from 16MB to 32MB using a file block size of 4MB, and using 16 nodes to process the read and write operations; slicing the part of the file from 32MB to 64MB using a file block size of 8MB, and using 32 nodes to process the read and write operations; slicing the part of the file after 64MB using a file block size of 16MB, and using all the remaining nodes to process the read and write operations.

[0007] Optionally, the initialization of the file includes calling a preset ID generation method generateID() in a specified configuration file config.hpp to initialize the file according to the configured progressive data striping configuration, slicing it according to different file block sizes, and using an array to record the ID of each generated file block.

[0008] Optionally, the execution steps of the ID generation method generateID() include: initializing the file according to the configured progressive data striping configuration, slicing it according to different file block sizes, calculating the file block start number chunkid of each file block of the striping component according to the start position and strip size of the striping component corresponding to each part of the file, and recording the file block number chunkid of each file block into the data PFLcomponents and returning it.

[0009] Optionally, the operation of the file system on the striped file includes: forwarding the write request of the prepared file block to the file system server by the file system client executing the file write operation, and parallel writing the corresponding file block into the file system by the file system server after receiving the request.

[0010] Optionally, forwarding the write request of the prepared file block to the file system server by the file system client executing the file write operation includes: the file system client initializes the metadata for the newly created file, where the metadata includes the file name, file path, permissions, owner information, and striping configuration, and the striping configuration includes, for example, the strip size and the number of stripes, and assigns a unique identifier to the file; the file system client calls a hash function to determine the storage location of each prepared file block, calculates the overflow part of the given offset within the current file block size, checks whether the given offset is aligned to the boundary of the file block size by adjusting the offset and the block size, calculates and returns the file block index corresponding to the given offset and block size, adjusts the offset and the block size according to the configuration of the striping component, and uses this information to calculate the index; the file system client calls the forward_write() function to initiate a file block write request to the file system server, and after receiving the request, the file system server stores the file block in the file system; the execution of the forward_write() function includes: S101, determine whether the size of the file is greater than 0. If not, generate an error prompt, end and exit; otherwise, jump to step S102; S102, calculate the boundary and number of the file block starting number chunkid of the current file block; S103, determine whether the current file block is the last file block. If it is the last file block, jump to step S104; otherwise, jump to step S105; S104, record the file block starting number chunkid of the current file block, update the file block starting numbers chunkid of the first and last file blocks, and expose the user buffer as the RDMA data source; S105, send an RPC request to the file system server to enable the file system server to process the RPC request, and monitor whether all the file block data of the file is written to the file system server. If all the file block data of the file is written, end and exit; otherwise, generate an error prompt, end and exit.

[0011] Optionally, when the file system server writes the corresponding file blocks into the file system in parallel after receiving a request, it includes that the file system server first verifies the file block offset. If the verification of the file block offset passes, it calls the RPC remote write function rpc_srv_write() to store the corresponding file block into the file system; the execution of the RPC remote write function rpc_srv_write() includes: S201, set RPC information; allocate space for the bulk transfer buffer; S202, process the file block start number chunkid of the file block passed in by the file system client through the RPC request; S203, determine whether there are still file blocks to be processed. If there are still file blocks to be processed, jump to step S204; otherwise, jump to step S208; S204, determine whether the file block is processed by the corresponding host host. If it is processed by the corresponding host host, jump to step S205; otherwise, jump to step S203; S205, dynamically adjust the size of the current file block; S206, determine whether the file block is the first or the last data block. If it is the first or the last data block, perform offset processing; S207, transfer the file block and start an asynchronous write operation; S208, wait for the task response, and send the result to the file system client after receiving the task response, and end and exit.

[0012] In addition, the present invention also provides a progressive data striping system for a task-level burst cache file system, including a microprocessor and a memory connected to each other. The microprocessor is programmed or configured to execute the progressive data striping method for the task-level burst cache file system.

[0013] In addition, the present invention also provides a computer-readable storage medium, in which a computer program or instruction is stored. The computer program or instruction is programmed or configured to execute the progressive data striping method for the task-level burst cache file system through a processor.

[0014] In addition, the present invention also provides a computer program product, including a computer program or instruction. The computer program or instruction is programmed or configured to execute the progressive data striping method for the task-level burst cache file system through a processor.

[0015] Compared with the prior art, the present invention can mainly achieve the following beneficial effects: The slice size is an important factor affecting I / O performance because it directly affects parallelism, cache hit rate, and the management overhead of the file system. Reasonably selecting the slice size can balance performance and resource usage, improving the throughput and response speed of the system. Excessive slices may lead to resource waste, while overly small slices may increase the management and transmission overhead. By optimizing the slice size according to different application scenarios and data access patterns, the I / O performance of the system can be significantly improved. In response to the problems of the prior art, the progressive data striping method for the task-level burst cache file system of the present invention stripe files in a differentiated manner, adopting different slice sizes for different parts of the file and configuring different numbers of storage nodes, reducing unnecessary I / O requests and latency, improving data storage performance and access efficiency, and further enhancing the throughput and response speed of the system. The present invention uses a progressive data striping strategy to perform progressive data striping and splitting on files, without splitting files using the same block size. Additionally, different numbers of nodes are allocated to process different striping components, which can effectively reduce the I / O load of the system and improve the parallel read and write performance of the system in application scenarios such as large-scale concurrent access, large file storage, and high-performance computing. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a schematic diagram of the basic process of the method according to an embodiment of the present invention.

[0017] Figure 2 It is a schematic diagram of progressive data striping in an embodiment of the present invention.

[0018] Figure 3 It is a schematic diagram of a striping component in an embodiment of the present invention.

[0019] Figure 4 It is a schematic diagram of the basic process of the forward_write() function in an embodiment of the present invention.

[0020] Figure 5 It is a schematic diagram of the basic process of the rpc_srv_write() function in an embodiment of the present invention.

[0021] Figure 6 It is a schematic diagram of the performance comparison of writing in an embodiment of the present invention.

[0022] Figure 7 It is a schematic diagram of the performance comparison of reading in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] To enable those skilled in the art to better understand the technical solution of the present invention, the following will further elaborate on the technical solution of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention.

[0024] As Figure 1 shown, the progressive data striping method for the task-level burst cache file system in this embodiment includes the following steps: configuring a striping component to perform progressive data striping on a file, initializing the file, and the file system operating on the striped file; the progressive data striping means that different parts of different file sizes are respectively segmented using different file block sizes and configured with different numbers of nodes for read and write operations, and starting from the head of the file, the size of each part monotonically increases, the file block size used monotonically increases, and the number of nodes used for read and write operations also monotonically increases. In high-performance computing scenarios, a large number of file read and write operations are involved. When processing file read and write operations, the slice size of the file is an important factor affecting data storage and access performance, as well as system I / O performance. When the slice size is too small, the concurrency performance can be improved, but each time a smaller slice is accessed, the system must perform additional metadata processing. Especially in a distributed file system, each slice requires a separate I / O request and metadata lookup. Therefore, if the slice is too small, the frequency of I / O requests will increase significantly. In addition, small slices mean that more blocks need to be accessed, which may result in more frequent access to different storage devices or storage blocks during each read or write, which may increase the addressing overhead of the storage device and reduce the overall performance. If the slice size is increased too much, the number of I / O requests can be reduced, but each large slice may occupy the resources of multiple storage nodes, resulting in overloading of some nodes while other nodes are idle, unable to fully utilize the parallel processing capabilities of the cluster. In addition, if the slice is too large and the actual amount of data to be accessed is small, it will waste the bandwidth of the storage system. Therefore, in order to optimize the file read and write performance in high-concurrency situations, the method of this embodiment performs progressive data striping processing on the file using a differential method, adopts different slice sizes for different parts of the file, and configures different numbers of storage nodes at the same time, reducing unnecessary I / O requests and delays, improving data storage performance and access efficiency, and further enhancing the throughput and response speed of the system.

[0025] As Figure 2As shown, when progressive data striping is performed in this embodiment, it includes splitting the 0 to 4MB part of the file using a file block size of 512KB, and using 1 node to handle the read and write operations; splitting the 4MB to 8MB part of the file using a file block size of 1MB, and using 3 nodes to handle the read and write operations; splitting the 8MB to 16MB part of the file using a file block size of 2MB, and using 8 nodes to handle the read and write operations; splitting the 16MB to 32MB part of the file using a file block size of 4MB, and using 16 nodes to handle the read and write operations; splitting the 32MB to 64MB part of the file using a file block size of 8MB, and using 32 nodes to handle the read and write operations; splitting the part after 64MB of the file using a file block size of 16MB, and using all the remaining nodes (62 in this embodiment) to handle the read and write operations, thereby forming a progressive layout of the file (Progressive File Layout, PFL). As Figure 3 As shown, there are six striping components in this embodiment: the first striping component contains the first 4MB of the file and only uses 1 node to process the file, the second striping component contains the range of 4MB to 8MB of the file and uses 3 nodes, the third striping component contains the range of 8MB to 16MB of the file and uses 8 nodes, until the last striping component uses all the nodes.

[0026] In this embodiment, initializing the file includes calling a preset ID generation method generateID() in the specified configuration file config.hpp to initialize the file according to the configured progressive data striping configuration, splitting it according to different file block sizes, and using an array to record the ID of each generated file block.

[0027] In this embodiment, the execution steps of the ID generation method generateID() include: initializing the file according to the configured progressive data striping configuration, splitting it according to different file block sizes, calculating the file block start number chunkid of each file block of the striping component based on the start position and strip size of the striping component corresponding to each part of the file, and recording the file block number chunkid of each file block into the data PFLcomponents and returning. Specifically in this embodiment, then the return value is passed to the parameter variable PFLchunkID, and the parameter variable PFLchunkID records the file block start ID (i.e., the file block start number chunkid) of each striping component.

[0028] In this embodiment, the operations of the file system on the striped file include: forwarding the write request of the prepared file block to the file system server through the file system client, and the file system server writes the corresponding file block into the file system in parallel after receiving the request.

[0029] In this embodiment, forwarding the write request of the prepared file block to the file system server through the file system client includes: the file system client initializes the metadata for the newly created file, and the metadata includes file name, file path, permissions, owner information, and striping configuration. The striping configuration includes strip size and number of stripes, etc., and assigns a unique identifier to the file; the file system client calls the hash function to determine the storage location of each prepared file block, calculates the overflow part of the given offset within the current file block size, checks whether the given offset is aligned to the boundary of the file block size by adjusting the offset and block size, calculates and returns the file block index corresponding to the given offset and block size, adjusts the offset and block size according to the configuration of the striping component, and uses this information to calculate the index; the file system client calls the forward_write() function to initiate a file block write request to the file system server, and the file system server stores the file block in the file system after receiving the request; as Figure 5 shown, the execution of the forward_write() function in this embodiment includes: S101, determine whether the size of the file is greater than 0. If not, generate an error prompt, end and exit; otherwise, jump to step S102; S102, calculate the boundary and number of the file block start number chunkid of the current file block; S103, determine whether the current file block is the last file block. If it is the last file block, jump to step S104; otherwise, jump to step S105; S104, record the file block start number chunkid of the current file block, update the file block start numbers chunkid of the first and last file blocks, and expose the user buffer as the RDMA data source; S105, send an RPC request to the file system server to make the file system server process the RPC request, and monitor whether all the file block data of the file is written to the file system server. If all the file block data of the file is written, end and exit; otherwise, generate an error prompt, end and exit.

[0030] In this embodiment, when the file system server writes the corresponding file blocks into the file system in parallel after receiving a request, it includes that the file system server first verifies the file block offset. If the verification of the file block offset passes, it calls the RPC remote write function rpc_srv_write() to store the corresponding file block into the file system; as Figure 4 shown, the execution of the RPC remote write function rpc_srv_write() in this embodiment includes: S201, set RPC (Remote Procedure Call, used for communication between the client and the storage server) information. Configuring the RPC information means setting the parameters related to the client and the storage server; allocate space for the bulk transfer buffer; S202, process the file block start number chunkid of the file block passed in by the file system client through the RPC request; S203, determine whether there are still file blocks to be processed. If there are still file blocks to be processed, jump to step S204; otherwise, jump to step S208; S204, determine whether the file block is processed by the corresponding host host. If it is processed by the corresponding host host, jump to step S205; otherwise, jump to step S203. In this embodiment, by determining whether the file block is processed by the corresponding host host, the temporary file system used realizes the decentralized distribution of file blocks to different hosts and allows them to process the corresponding data blocks respectively, reducing the single-point bottleneck and improving the I / O performance; S205, dynamically adjust the size of the current file block. Oversized slices may cause resource waste, and undersized slices may cause increased management and transmission overhead. By optimizing the slice size according to different application scenarios and data access patterns, the I / O performance of the system can be significantly improved, that is, adjusted according to the previous Progressive File Layout (PFL); S206, determine whether the file block is the first or the last data block. If it is the first or the last data block, perform offset processing; by judging whether the number of the current data block is equal to the start block number or the end block number, to determine whether it is the first or the last data block. If it is the first data block, the excess part in the file offset needs to be subtracted; if it is the last data block, the part less than a complete file block at the end of the file needs to be subtracted to ensure that the data can be correctly aligned and stored; S207, transmit the file block and start an asynchronous write operation; S208, wait for the task response, and send the result to the file system client after receiving the task response, and end and exit.

[0031] The file system client calls the forward_write() method to perform a file write operation, forwarding the prepared file block write request to the file system server. After receiving the request, the file system server calls the rpc_srv_write() method to write the corresponding file blocks to the file system in parallel. First, the file system initializes the metadata for the newly created file. The metadata includes the file name, file path, permissions, owner information, and striping configuration (such as stripe size and number of stripes), and assigns a unique identifier to the file. Then, the forward_write() method is called to initiate a request to the file system server. After receiving the request, the file system server calls the rpc_srv_write() method to store the file blocks requested by the client in the file system.

[0032] To verify the progressive data striping method for the task-level burst buffer file system in this embodiment, IOR is used for testing in this embodiment. IOR (Input / Output Review) is a widely used I / O performance testing tool, usually used to evaluate the throughput and performance of the storage subsystem in a high-performance computing system. By simulating high-load I / O scenarios of parallel file systems (such as Lustre, GPFS, BeeGFS, etc.), the performance of the storage system is tested and analyzed. To evaluate the read and write performance of progressive data striping in a high-performance computing environment, the IOR test tool is used to perform read and write performance tests and compare with traditional file systems using non-progressive data striping. The test file sizes are gradually increased from 128MB, 256MB, 512MB, 1024MB to 2048MB, and the average value is taken after each test is run 10 times. The test results are shown in Figure 6 and Figure 7 , Figure 6 show the comparative experimental results of the IOR test read operation, Figure 7 show the comparative experimental results of the IOR test write operation. The units of the test results are all MB / S. See Figure 6 and Figure 7It can be seen that progressive data striping performs better in both writing and reading performance. Especially when dealing with large data blocks (such as 2048MB), the performance improvement is the most obvious. This indicates that the progressive data striping method can more effectively utilize storage resources and optimize performance by dynamically adjusting the strip size. In the case of different block sizes, the performance growth of progressive data striping is relatively large, while non-progressive data striping shows relatively stable and inefficient performance. Especially in the case of smaller block sizes, it fails to effectively improve the bandwidth utilization rate. Generally speaking, progressive data striping can provide higher throughput and lower latency under different read-write loads, and is suitable for application scenarios that need to process a large amount of data and high concurrency. Non-progressive data striping is more suitable for processing scenarios with uniform loads, but in the case of complex tasks or large amounts of data, the performance improvement is limited.

[0033] In summary, the progressive data striping adopted in the progressive data striping method for the task-level burst cache file system in this embodiment is a method that optimizes the file system performance by gradually adjusting the file striping strategy according to the file layout. In a high-concurrency scenario, the striping of files directly affects the way data is stored and accessed. A smaller striping unit will significantly increase the number of I / O requests, while a larger striping unit, although it will reduce the number of I / O requests, may cause greater latency because each I / O request is larger and takes longer to complete. In addition, a larger strip may bring excessive data reading / writing, resulting in waste of disk bandwidth. Therefore, using the progressive data striping strategy to allocate strips of different sizes to different parts of the file, this progressive adjustment strategy can appropriately split the file, thereby improving I / O efficiency, reducing latency, and further enhancing the file read-write performance in a high-concurrency environment.

[0034] In addition, this embodiment also provides a progressive data striping system for a task-level burst cache file system, including a microprocessor and a memory connected to each other. The microprocessor is programmed or configured to execute the progressive data striping method for the task-level burst cache file system.

[0035] In addition, this embodiment also provides a computer-readable storage medium, in which a computer program or instruction is stored. The computer program or instruction is programmed or configured to execute the progressive data striping method for the task-level burst cache file system through a processor.

[0036] In addition, this embodiment also provides a computer program product, including a computer program or instruction. The computer program or instruction is programmed or configured to execute the progressive data striping method for the task-level burst cache file system through a processor.

[0037] Those skilled in the art should understand that the technical solutions provided by the present invention can be in the form of a method, a system, or a computer program product. Therefore, the present invention can be implemented in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can be in the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code. The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in the process Figure 1 one process or multiple processes and / or blocks Figure 1 These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in the process Figure 1 one process or multiple processes and / or blocks Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0038] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements should also be regarded as within the protection scope of the present invention.

Claims

1. A progressive data striping method for a task-level burst cache file system, characterized in that: The method comprises the following steps: configuring a striping component to perform progressive data striping on a file, initializing the file, and the file system operating the striped file; the progressive data striping means that different parts of the file with different sizes are divided using different file block sizes and different numbers of nodes are configured for read and write operations, and starting from the head of the file, the size of each part increases monotonically, the file block size used increases monotonically, and the number of nodes used for read and write operations also increases monotonically.

2. The progressive data striping method for a task-level burst cache file system according to claim 1, characterized in that: The progressive data striping includes segmenting the portion of the file from 0 to 4 MB using a file block size of 512 KB, and processing the read and write operations using one node; The 4MB to 8MB portion of the file is split using a 1MB file block size, and read and write operations are processed using 3 nodes; the 8MB to 16MB portion of the file is split using a 2MB file block size, and read and write operations are processed using 8 nodes; the 16MB to 32MB portion of the file is split using a 4MB file block size, and read and write operations are processed using 16 nodes; the 32MB to 64MB portion of the file is split using an 8MB file block size, and read and write operations are processed using 32 nodes; the portion after 64MB of the file is split using a 16MB file block size, and read and write operations are processed using all remaining nodes.

3. The progressive data striping method for a task-level burst cache file system according to claim 1, characterized in that: The initialization file includes calling the preset ID generation method generateID() in the specified configuration file config.hpp to initialize the file according to the configured progressive data striping configuration, split the file into blocks of different sizes, and use an array to record the ID of each generated file block.

4. The progressive data striping method for a task-level burst cache file system according to claim 3, characterized in that: The execution steps of the ID generation method generateID() include: initializing the file according to the configured progressive data striping configuration, dividing the file according to file block sizes of different sizes, calculating the file block starting number chunkid of each file block of the striping component according to the starting position and stripe size of the striping component corresponding to each part of the file, and recording the file block number chunkd of each file block in the data PFLcomponents and returning it.

5. The progressive data striping method for a task-level burst cache file system according to claim 1, characterized in that: The file system operates the striped files by: executing a file write operation through the file system client to forward a prepared file block write request to the file system server, and the file system server writes the corresponding file blocks into the file system in parallel after receiving the request.

6. The progressive data striping method for a task-level burst cache file system according to claim 5, characterized in that: The method of executing a file write operation through a file system client to forward a prepared file block write request to a file system server includes: the file system client initializes metadata for a newly created file, the metadata includes a file name, a file path, permissions, owner information, and a striping configuration, the striping configuration includes, for example, a stripe size and a number of stripes, and assigns a unique identifier to the file; the file system client calls a hash function to determine the storage location of each prepared file block, calculates an overflow portion of a given offset within a current file block size, checks whether a given offset has been aligned to a file block size boundary by adjusting the offset and the block size, calculates and returns a file block index corresponding to a given offset and a block size, adjusts the offset and the block size according to the configuration of the striping component, and uses this information to calculate an index; the file system client calls a forward write function forward_write() to initiate a file block write request to the file system server, and the file system server stores the file block in the file system after receiving the request; the execution of the forward write function forward_write() includes: S101, determine whether the file size is greater than 0, if not, generate an error prompt, end and exit; otherwise, jump to step S102; S102, calculating the boundary and number of the file block starting number chunkid of the current file block; S103, determining whether the current file block is the last file block, if it is the last file block, jumping to step S104, otherwise, jumping to step S105; S104, recording the file block starting number chunkid of the current file block, updating the file block starting numbers chunkid of the first and last file blocks, and exposing the user buffer as an RDMA data source; S105, sending an RPC request to the file system server so that the file system server processes the RPC request and monitors whether all the file block data of the file is written to the file system server. If all the file block data of the file is written, the process ends and exits; otherwise, an error prompt is generated, the process ends and exits.

7. The progressive data striping method for a task-level burst cache file system according to claim 5, characterized in that: When the file system server writes the corresponding file block to the file system in parallel after receiving the request, the file system server first verifies the file block offset, and if the verification of the file block offset is passed, calls the RPC remote write function rpc_srv_write() to store the corresponding file block in the file system; the execution of the RPC remote write function rpc_srv_write() includes: S201, setting RPC information; allocating space for batch transmission buffer; S202, processing the file block starting number chunkid of the file block input by the file system client through the RPC request; S203, determine whether there are still file blocks to be processed, if there are still file blocks to be processed, jump to step S204; otherwise, jump to step S208; S204, determining whether the file block is processed by the corresponding host host, if so, jump to step S205; otherwise, jump to step S203; S205, dynamically adjust the size of the current file block; S206, determining whether the file block is the first or last data block, and if it is the first or last data block, performing offset processing; S207, transmitting the file block and starting an asynchronous write operation; S208, waiting for the task response, and sending the result to the file system client after receiving the task response, ending and exiting.

8. A progressive data striping system for a task-level burst cache file system, comprising a microprocessor and a memory connected to each other, characterized in that: The microprocessor is programmed or configured to execute the progressive data striping method for a task-level burst cache file system as claimed in any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program or instruction stored therein, characterized in that: The computer program or instruction is programmed or configured to execute the progressive data striping method for a task-level burst cache file system as claimed in any one of claims 1 to 7 through a processor.

10. A computer program product comprising a computer program or instructions, characterized in that The computer program or instruction is programmed or configured to execute the progressive data striping method for a task-level burst cache file system as claimed in any one of claims 1 to 7 through a processor.

Citation Information

Patent Citations

  • Cache file system communication method and system based on MP and RDMA

    CN111416872A

  • Low-latency file system address space management method and system and medium

    CN111522507A

  • Self-adaptive rapid increment pre-reading method for wide area network file system

    CN111787062A

  • Distributed file system-oriented high-concurrency read-write optimization system, medium and equipment

    CN116737685A

  • Method, device, and computer program product for data management

    US20210342334A1