Asynchronous checkpoint cache control method and device for high-performance computing system and medium
By dividing the HPC system into SBB groups and PFS groups, the checkpoint caching and refresh process is optimized, solving the problems of insufficient cache space and low storage utilization efficiency, achieving efficient checkpoint caching and refresh, and improving the execution efficiency of computing tasks.
Patent Information
- Application Number
- CN202511505009.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-21
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-10-21
AI Technical Summary
In existing HPC systems, asynchronous checkpointing technology suffers from performance bottlenecks in the checkpoint caching process due to insufficient cache space and low efficiency in utilizing new storage technologies, thus affecting the efficiency of computing tasks.
The compute nodes of the HPC system are divided into SBB groups and PFS groups. Checkpoint caching and flushing are performed using client-server mode and file I/O mode, respectively. The compute node partitioning method of SBB group and PFS group is optimized to minimize write time difference. RDMA technology is used to interact directly with SBB server to reduce data transfer between compute nodes.
It improves checkpoint caching and refresh efficiency, reduces the blocking time of computing tasks, makes full use of the heterogeneous storage resources of HPC systems, and avoids the performance bottleneck of a single storage medium.
Smart Images

Figure CN120973558A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of high performance computing (HPC) systems, and particularly relates to an asynchronous checkpoint cache control method for a high performance computing system, a device and a medium. BACKGROUND
[0002] With the continuous expansion of the scale of HPC systems, the number of processor cores and accelerators integrated in the system is increasing, which correspondingly increases the probability of hardware failure, and the mean time between failures (MTBF) of the system presents a significant downward trend. The mean failure-free time of an Exascale System in the future may be measured in minutes. Scientific computing applications in HPC systems usually need to perform large-scale numerical simulation or handle complex modeling problems, and the execution time of a single task may be tens to hundreds of hours. During the execution of the task, there is a risk of job interruption due to system failure. To cope with the risk of job interruption due to system failure, checkpointing technology is widely used as a fault-tolerant mechanism for HPC systems in the prior art.
[0003] Checkpointing technology usually directly writes checkpoints into a parallel file system (PFS) for storage. However, the size of a checkpoint for a large-scale application may reach hundreds of GB or even tens of TB. In this scenario, the limited bandwidth of the PFS will become a performance bottleneck, and the PFS will have difficulty completing the writing of the checkpoint within an acceptable time, which will significantly affect the computing efficiency of the application. Asynchronous checkpointing technology can alleviate the above problems. Asynchronous checkpointing technology first caches the checkpoint in local storage, and then a background process flushes the data cached in the local storage to the PFS, thereby improving the reliability of the checkpoint. Since asynchronous checkpointing technology is used, the application only needs to be blocked when the checkpoint is written to the cache, and the data flushing operation in the background is transparent to the application, so the application does not need to be blocked, and since the writing speed of the cache is usually very high, the application using asynchronous checkpointing technology can quickly remove the blocking caused by the checkpoint.
[0004] However, with the continuous growth of the computing power of HPC systems, the computing scale of the application allowed by the system also increases, and the size of its checkpoint also increases accordingly. However, the capacity of the local storage of the computing node in the HPC system has not increased correspondingly, which causes certain limitations when the application uses asynchronous checkpointing technology: 1. Insufficient cache space. Since the asynchronous checkpointing technology usually uses local memory as cache, however, large-scale applications usually use a large amount of memory during runtime, and the remaining memory capacity available for checkpoint cache is limited, so that the asynchronous checkpointing technology cannot complete the checkpoint cache process due to insufficient cache space, in which case the asynchronous checkpointing technology can only degenerate into synchronous checkpointing technology, that is, directly writing the checkpoint to PFS, and due to the limited bandwidth of PFS, the synchronous checkpointing technology usually needs to block the application for a long time.
[0005] 2. Low utilization efficiency of new storage. The current HPC system gradually introduces a new type of storage such as shared burst buffer (SBB), which usually uses high-performance NVMe SSD as storage device. Compared with local memory, this type of device usually has larger capacity, and compared with PFS, it has higher bandwidth, but compared with PFS, SBB has lower reliability due to limited SSD life and other factors. To maximize the reliability of the checkpoint, if the asynchronous checkpointing technology needs to use SBB cache, it needs to flush the checkpoint in SBB to PFS. However, the mainstream implementation library of the current asynchronous checkpointing technology relies on the distributed file system to write and read data from SBB. The distributed file system is essentially a unified organization and abstraction of SBB, and the underlying details of SBB cannot be controlled by the caller of the file system interface. When reading and writing SBB through the distributed file system, for the checkpoint cached in SBB, the mainstream implementation library of the current asynchronous checkpointing technology cannot make SBB directly write the cached checkpoint to PFS, but needs the computing node to read the checkpoint from the local buffer through the distributed file system, and then write the checkpoint in the local buffer of the computing node to PFS. In the above process, the checkpoint cached in SBB needs to be transferred by the computing node to flush to PFS. In large-scale application scenarios, the checkpoint is usually large, and the data transferred by the computing node is large, which will lead to low efficiency of the process of flushing the cached checkpoint to PFS. SUMMARY
[0006] The technical problem to be solved by the present application is that the existing technology has the above-mentioned problems. The present application provides an asynchronous checkpoint cache control method, device and medium for a high-performance computing system, which has the advantages of simple implementation method, low cost, high storage utilization efficiency, high cache and flushing efficiency, can effectively alleviate the performance bottleneck in the checkpoint process of large-scale applications in HPC systems, improve the checkpoint efficiency, and reduce the impact on computing tasks.
[0007] To solve the above technical problems, the technical scheme provided by the present application is as follows: An asynchronous checkpoint cache control method of a high performance computing (HPC) system, comprising the steps of: dividing computing nodes in the HPC system into SBB groups and PFS groups, wherein the SBB groups use SBBs as checkpoint caches and the PFS groups use PFSs as checkpoint caches, and determining an optimal computing node division manner for the SBB groups and the PFS groups to complete checkpoint writing in a minimum time difference; starting a server on each SBB node and embedding a client in each process on each computing node in the SBB groups, each client communicating with a SBB server; in a checkpoint cache phase, using different modes to write checkpoints of the SBB groups and the PFS groups to SBBs and PFSs respectively, wherein for the SBB groups, a client-server mode is used to write checkpoints to SBBs, and for the PFS groups, a file I / O mode is used to write checkpoints to PFSs; in a checkpoint flush phase, flushing the checkpoints in the cache to PFSs, wherein for the checkpoints already cached in PFSs, no change is made, and for the checkpoints cached in SBBs, a client-server mode is used to flush the checkpoints to PFSs.
[0008] Further, the determining of the optimal computing node division manner for the SBB groups and the PFS groups to complete checkpoint writing in a minimum time difference comprises: constructing the following computing node division model: F( )=∣ ∣ wherein, is the number of nodes concurrently writing to PFSs, is the number of nodes concurrently writing to SBBs, is the total amount of data written to SBBs, is the total amount of data written to PFSs, , are the bandwidths of SBBs and PFSs respectively when the concurrency is and respectively, represents the total writing time of the SBB groups, represents the total writing time of the PFS groups, F( ) is a function of and , and represents the absolute difference between the writing times of the SBB groups and the PFS groups; by finding the solution that makes F( ) minimum, the optimal computing node division manner for the SBB groups and the PFS groups to complete checkpoint writing in a minimum time difference is obtained.
[0009] Further, the constraint condition of setting the following is also included:
[0010]
[0011]
[0012]
[0013]
[0014]
[0015] wherein, D is the total amount of checkpoint data of the job, N is the number of nodes used for running the job, is a function of , and the function relationship is represented by , is a function of , and the function relationship is represented by .
[0016] In the process of finding the solution that makes F( ) take the minimum value, by obtaining the bandwidth of SBB and PFS under different concurrency, the concurrency-SBB bandwidth table and the concurrency-PFS bandwidth table are formed, the concurrency-SBB bandwidth table and the concurrency-PFS bandwidth table are traversed, all possible bandwidth values satisfying the constraint condition are found out and the corresponding F( ) values are calculated, and the values of and corresponding to the time when F( ) takes the minimum value among the multiple F(
[0017] ) values found out are taken as the number of computing nodes of the SBB group and the PFS group respectively. Further, in the checkpoint caching stage, the use of the client-server mode to write the checkpoint to the SBB includes: The client first checks whether there is an unfinished checkpoint flushing operation, and if there is, waits for the completion of the checkpoint flushing operation, otherwise starts the current checkpoint caching process; After starting the current checkpoint caching process, the client registers each state that needs to be cached by the process as an RDMA memory area; The client sends a checkpoint caching RPC request to the corresponding server for each state to be cached, and the request carries the start address and length information of the current state data in the process memory; The client is blocked so that the current to-be-cached state is not modified until a positive reply to the RPC request is received from the server; After the client receives a positive reply from the server regarding all states, it determines that the SBB checkpointing is complete, unblocks the client, and continues to perform the computing task.
[0018] Further, the server sends a positive reply to the RPC request sent by the client, including: The server pre-creates a buffer of a fixed size and registers it as an RDMA memory region in response to receiving the RPC request sent by the client regarding checkpointing; The server initializes an RDMA operation counter cnt to represent the number of operations required to complete the caching of one state of the current client; The server uses RDMA to bypass the operating system kernel to read the state from the memory of the corresponding client process according to the start address and length of the current state into the RDMA memory region, and writes the state in the RDMA memory region into the local storage device, and updates the value of the RDMA operation counter cnt; It is determined whether the value of the RDMA operation counter reaches a preset value. If not, the server uses RDMA to bypass the operating system kernel to read the state from the memory of the corresponding client process according to the start address and length of the current state into the RDMA memory region, and writes the state in the RDMA memory region into the local storage device, and updates the value of the RDMA operation counter cnt; if yes, it is determined that all state data of the current RPC request has been read and written into the local storage device, and the server sends a positive reply to the corresponding client RPC request.
[0019] Further, in the checkpoint flush phase, the data cached in the SBB is flushed to the PFS using the client-server mode, including: Each process on each computing node of the SBB group sends a checkpoint flush RPC request to the corresponding server; After the client sends the checkpoint flush RPC request to the server, it returns and continues to perform the computing task; After the server receives the checkpoint flush request, it creates a buffer and reads the corresponding checkpoint cache of the client in the local storage device into the buffer, and then uses the file system interface provided by the PFS to flush the data in the buffer to the PFS; After the server flushes all checkpoint caches of a client to the PFS, it sends a positive reply to the corresponding client RPC request; The client checks whether the server sends a positive reply to the RPC request to determine whether the last checkpoint flush is complete when performing the next checkpoint caching.
[0020] Further, in the checkpoint flush phase, it also includes keeping the connection established by each client with the SBB server in the checkpoint cache phase, so that the client uses the service from the same server in the two phases of checkpoint cache and flush; In the checkpoint flush phase, it also includes that the PFS group and the computing nodes of the SBB group keep synchronization when performing the computing task, that is, when the computing nodes of the SBB group send the RPC request of checkpoint flush to the server, the computing nodes of the PFS group block the computing, and when the computing nodes of the SBB group continue to perform the computing task, the computing nodes of the PFS group unblock to continue the computing.
[0021] Further, it also includes that in the restart phase, reading the checkpoint of all application processes from the PFS to the corresponding area in the process address space of the specified state to restore the computing of the application process before the failure.
[0022] An electronic device, comprising a processor and a memory, the memory is used to store a computer program, and the processor is used to execute the computer program to perform the method as described above.
[0023] A computer readable storage medium storing a computer program, the computer program is executed by a processor to implement the method as described above.
[0024] Compared with the prior art, the beneficial effects of the present application are: 1. The present application fuses the PFS and SBB two kinds of heterogeneous storage in the HPC system as the cache of the checkpoint, fully utilizes the SBB and PFS two kinds of heterogeneous storage resources in the HPC system to cache the checkpoint, can disperse the I / O pressure, avoids the performance bottleneck that may exist in the single storage, and shortens the time spent for checkpoint cache.
[0025] 2. The present application optimizes the flush of the SBB cache, adopts the "client-server" mode to interact with the remote heterogeneous storage SBB, instead of accessing the SBB through the distributed file system, when the checkpoint cached in the SBB needs to be flushed to the PFS, the computing node sends the request as the client, the SBB receives the request as the server, and directly flushes the data to the PFS, reduces the participation of the computing node in the checkpoint flush process, does not need the computing node to transfer the data, significantly improves the checkpoint flush efficiency, and reduces the influence on the computing task. BRIEF DESCRIPTION OF DRAWINGS
[0026] Figure 1 is the implementation flow diagram of the asynchronous checkpoint cache control method of the high-performance computing system of the present embodiment. DETAILED DESCRIPTION
[0027] The application is further described below in conjunction with the accompanying drawings and specific preferred embodiments, but it will be apparent to those skilled in the art that the application is not limited to the specific embodiments described herein.
[0028] For the convenience of understanding, first exemplary background of related art involved in the present application is introduced.
[0029] In the field of HPC, application programs usually perform large-scale numerical simulation tasks or complex modeling tasks, and a single running cycle can last for tens or even hundreds of hours. In recent years, the scale of HPC systems has been expanding, and the MTBF has shown a trend of shortening. In particular, in future E-class systems, the MTBF can be only in minutes, and checkpoint technology is a fault-tolerant mechanism to cope with the risk of job interruption caused by system-level failures. The core mechanism of checkpoint technology is to persistently save the key state of the application process at a specific execution time. The application process usually performs periodic checkpoint operations. In order to ensure the consistency of the key state to be saved and prevent the key state from changing due to the continued execution of the application process during the saving process, the application process is usually required to be in a blocked state during the checkpoint execution.
[0030] When encountering system failure, the application process loads the available latest checkpoint version from the persistent storage to recover its key state. For application programs using checkpoint technology, when they encounter system failure, they do not have to restart from the beginning, but can recover from the checkpoint, which can greatly avoid the loss of computing progress. Checkpoint technology can be divided into traditional checkpoint technology and asynchronous checkpoint technology according to its I / O characteristics. Traditional checkpoint technology is also called synchronous checkpoint technology due to its synchronous write to PFS characteristics. Traditional checkpoint technology is the early checkpoint technology, which usually directly writes the checkpoint to PFS for storage. However, the bandwidth of PFS is often limited, and when a large-scale application process generates TB-level checkpoint data, synchronous writing to PFS will cause significant I / O bottlenecks, forcing the application process to block for a long time, which seriously affects the computing efficiency.
[0031] The asynchronous checkpointing technique can alleviate the problem of too long application process blocking time of the traditional checkpointing technique or the synchronous checkpointing technique. The asynchronous checkpointing technique first caches the checkpoint in the local storage, and then a background process flushes the data cached in the local storage to the PFS, so as to improve the reliability of the checkpoint. The PFS is a file system specially used for HPC systems, which is shared among multiple users, multiple jobs and multiple nodes. The PFS needs to provide not only high-concurrency data access capability, but also high-capacity data persistent storage capability. The capacity of the PFS can reach the PB level. The NVMe SSD is an electronic storage device without mechanical components. The NVMe SSD uses a flash memory chip to store data, and the data access does not involve the movement of mechanical components, so the NVMe SSD can quickly access data and has high bandwidth. However, the NVMe SSD is manufactured by using advanced semiconductor technology, and the cost is high. It is difficult to deploy a large number of NVMe SSDs to achieve the PB-level storage capacity required by the PFS. The HDD is a traditional mechanical storage device. The HDD accesses data by moving a read-write head and one or more rotating magnetic disks. The mechanical seeking process limits the bandwidth of the HDD, so the bandwidth of the HDD is lower than that of the NVMe SSD. However, the material cost and manufacturing cost of the HDD are lower than those of the NVMe SSD. Therefore, considering the cost, the underlying storage devices of the PFS are usually deployed in large quantities by using HDDs in the HPC system, so as to achieve the PB-level storage capacity required by the PFS in an economical way. Because the PFS uses the HDD as the storage device, the bandwidth of the PFS is limited. Therefore, when an application process writes data to the PFS in a large-scale and high-concurrency mode, such as writing TB-level checkpoint data, a long time is often required.
[0032] The SBB is a new type of storage. Conceptually, the SBB is a high-speed shared storage system, which is usually composed of multiple high-performance NVMe SSDs and connected to the computing nodes through a high-speed network such as InfiniBand or Omni-Path. Conceptually, the capacity of the SBB is usually smaller than that of the PFS and larger than that of the local memory of the computing node, and the theoretical peak bandwidth of the SBB is usually lower than that of the local memory of the computing node and higher than that of the PFS. The main design goal of the SBB is to alleviate the I / O bottleneck of the PFS, and to provide an I / O buffer layer between the high-speed local memory of the computing node and the low-speed PFS. The application process can cache data in the SBB, and then flush the data to the PFS in the background. The SBB is usually composed of a large number of nodes and storage devices, and the underlying physical structure of the SBB can be very complex. The current asynchronous checkpointing technique can only rely on the unified abstraction of the SBB performed by the distributed file system to access data of the SBB.
[0033] SBB and PFS are storage shared among multiple applications and multiple nodes in HPC system, and their capacity is usually sufficient, while their bandwidth is related to their concurrency, i.e., the number of computing nodes concurrently writing to the shared storage such as SBB or PFS. The following gives an example to illustrate the relationship between the bandwidth of shared storage and its concurrency, for example, when 2 computing nodes write to PFS at the same time, the bandwidth of PFS is x GB / s, and when the concurrency of writing to PFS is increased, i.e., 4 computing nodes write to PFS, the bandwidth of PFS is y GB / s, since the concurrency of the above two cases is at a low level, which is not sufficient to fully utilize the bandwidth of PFS, therefore x is not equal to y (when the above two cases can fully utilize the bandwidth of PFS, i.e., reach the peak bandwidth of PFS, x is equal to y), and since the concurrency of the second case is higher than that of the first case, y > x. The bandwidth of SBB also shows the same rule as PFS, and is also related to the concurrency.
[0034] Although SBB uses NVMe SSD as storage device, its theoretical peak bandwidth can be higher than that of PFS using HDD as storage device, however, when the concurrency of SBB is insufficient and the concurrency of PFS is high, the actual bandwidth of SBB may be lower than that of PFS. In addition, although in an ideal design, the SBB in the HPC system should have a peak bandwidth much larger than that of PFS, so as to effectively alleviate the I / O bottleneck between the computing nodes and the persistent storage layer. However, in actual deployment, there is a significant difference in the construction and operation and maintenance budget of HPC system. For budget-constrained systems, the number of NVMe SSDs they can deploy may not be sufficient to form a large-scale, high-bandwidth SBB cluster. In this case, when facing the high concurrency and high throughput data write demand generated by large-scale parallel applications, such as checkpointing demand, the limited-scale SBB may not be able to provide sufficient aggregated write bandwidth, thereby becoming a bottleneck of system performance, and failing to fully meet the demand of application for high-speed data caching and exchange.
[0035] Although the asynchronous checkpointing technology can reduce the blocking time of application processes, in the current HPC system, on the one hand, due to the limited capacity of local memory, the local memory cannot be used to cache the complete checkpoint of large-scale applications. With the expansion of the scale of HPC system, the computing scale of application program also increases, and when the large-scale application program runs, the local memory resource is largely occupied by the computation itself, and the remaining available local memory is insufficient to cache the complete checkpoint data, so that the asynchronous checkpointing technology can only degenerate into synchronous checkpointing technology, i.e., all checkpoint data is written to the slow PFS without any form of caching, causing the application process to be blocked for a long time, and affecting the execution of its computing task.
[0036] On the other hand, when using the new storage SBB with sufficient capacity as a cache, the way of relying on the distributed file system to flush the SBB cache is inefficient. Although HPC systems gradually deploy SBB, SBB is usually used as a storage device with NVMe SSD, and its bandwidth is usually higher than that of PFS in design concept. At the same time, SBB is a storage resource that can be shared between multiple jobs, and its capacity is usually sufficient for checkpoint storage. SBB is suitable for checkpoint cache because of its high bandwidth and large capacity. However, when using SBB as a cache, since the current asynchronous checkpoint technology relies on the distributed file system to access SBB, when flushing the SBB cache to the final storage PFS, the data must first be read by the computing node to its local, and then written to PFS by the computing node from the local. The above process of checkpoint cache flushing requires the computing node to transfer the data, and the checkpoint data volume of large-scale applications is huge. This transfer process significantly reduces the flushing efficiency and has a great interference to the application process of the computing node.
[0037] To solve the above problems, the application constructs a new framework HeteroCheck based on asynchronous checkpoint technology, which can effectively alleviate the performance bottleneck that may exist in the checkpoint process of large-scale applications in HPC systems, improve the checkpoint efficiency, and reduce the impact on computing tasks. At the same time, the application uses a client-server mode instead of a distributed file system in the asynchronous checkpoint technology, which enables the computing node to start the client and the SBB node to start the server. When flushing the checkpoint of the SBB cache, the client sends a flushing request to the server, and the server directly writes the data stored in the SBB to the PFS after receiving the request, without the computing node transferring the data. Therefore, the data flow during the checkpoint flushing of the SBB cache can be simplified from "SBB→computing node→PFS" to "SBB→PFS", which can greatly reduce the network overhead caused by data movement in high-performance computing systems, and also offload the SBB flushing task of the computing node, release the CPU and network resources of the computing node, and thus reduce the impact on the computing task of the computing node.
[0038] Compared with the strategy of storing checkpoints into a single storage medium, the heterogeneous storage strategy of the application can fully integrate the diversified storage resources in the HPC system, avoid the performance bottleneck that may be caused by a single medium, and thus significantly shorten the time consumption of checkpoint operation. Meanwhile, by using the "client-server" mode to interact with the SBB remote storage resource, when the checkpoint in the SBB needs to be refreshed to the PFS, the computing node will act as a client and the SBB will act as a server. The computing node only needs to send a checkpoint refresh request to the server on the SBB, and the data cached in the checkpoint will be directly refreshed to the PFS by the server. Compared with the traditional way of reading and writing the SBB through the distributed file system, since the computing node no longer needs to transfer data during the checkpoint refresh process, the efficiency of checkpoint refresh can be effectively improved.
[0039] For large-scale applications, the local memory of the computing node may not be sufficient to completely cache the checkpoints of the large-scale application. In this case, the current asynchronous checkpoint technology can only be degraded to the synchronous checkpoint technology, and all checkpoints are directly written to the slow PFS, which significantly increases the blocking time of the application process. The HeteroCheck framework of the application can use both PFS and SBB as caches, so that the SBB can share the pressure of the PFS and relieve the performance bottleneck that may occur on the PFS. At the same time, for large-scale applications, SBB alone may not be able to fully meet the high demand of the application process for aggregated bandwidth. The dual-storage caching strategy of the application can fully utilize the PFS to relieve the pressure of the SBB and avoid the performance bottleneck that may exist in a single storage.
[0040] For large-scale applications, the checkpoint thereof usually reaches the order of PB. When using SBB as a cache, the current asynchronous checkpoint technology relies on the distributed file system to access the data in the SBB. During the checkpoint refresh process, the distributed file system transports the data to the computing node, and then the computing node flushes the data into the PFS. The above checkpoint refresh process needs the computing node to participate as a transit node, which not only has the limitation of low efficiency, but also may interfere with the application process on the computing node. The HeteroCheck framework of the application can unload the task of refreshing the SBB checkpoint cache directly to the SBB server, without the participation of the computing node. This can reduce the participation of the computing node in the checkpoint refresh process, and the computing node does not need to transfer data, thereby effectively improving the refresh efficiency.
[0041] The application will be further described below in conjunction with specific embodiments.
[0042] As Figure 1As shown, the steps of the asynchronous checkpoint cache control method of the high-performance computing system of the embodiment include: Step S01. Divide the computing nodes in the HPC system into two groups, including an SBB group and a PFS group, wherein the SBB group uses SBB as the checkpoint cache, the PFS group uses PFS as the checkpoint cache, and the optimal computing node division manner of the SBB group and the PFS group is determined to make the SBB group and the PFS group complete checkpoint writing with the minimum time difference; Step S02. Start a server on each SBB node, and embed a client in each process on each computing node in the SBB group, each client communicates with an SBB server, so that the SBB provides direct services for the computing nodes; Step S03. In the checkpoint cache stage, use different modes to write the checkpoints of the SBB group and the PFS group to SBB and PFS respectively, wherein for the SBB group, use the client-server mode to write the checkpoints to SBB, and for the PFS group, use the file I / O mode to write the checkpoints to PFS; Step S04. In the checkpoint flush stage, flush the checkpoints in the cache to PFS, wherein the checkpoints already cached in PFS remain unchanged, and the checkpoints cached in SBB are flushed to PFS using the client-server mode.
[0043] In the embodiment, the asynchronous checkpoint framework HeteroCheck is constructed, two kinds of heterogeneous storage SBB and PFS in the HPC system are fused as the checkpoint cache by the HeteroCheck framework, the possible checkpoint cache performance bottleneck of single storage is avoided; at the same time, the “client-server” mode is adopted to interact with the remote heterogeneous storage SBB, when the checkpoints cached in SBB need to be flushed to PFS, the computing node sends a request as a client, the SBB receives the request as a server, and directly flushes the data to PFS, reducing the participation of the computing node in the checkpoint flush process, without the computing node transferring data, improving the flush efficiency.
[0044] In the embodiment, multiple storages are used to avoid the problem that a single storage cannot store the checkpoints completely due to limited capacity, and to avoid the problem that a single storage causes a long checkpoint time and a long application blocking time due to insufficient bandwidth. Specifically, the HeteroCheck framework uses two storages, SBB and PFS, to store the checkpoints, and the cached data of a checkpoint of an application is distributed in the two storages. The computing nodes are divided into two groups in the embodiment, one group stores the checkpoints in SBB, corresponding to an "SBB group", and the other group stores the checkpoints in PFS, corresponding to a "PFS group". The optimal number of nodes of the SBB group and the PFS group is determined by using an algorithm based on minimizing the tail latency effect, so that the two groups of nodes can complete the checkpoint writing at the same time as much as possible. The computing nodes with the smallest numbers are sequentially added to the SBB group until the required number of nodes of the SBB group is reached, and the remaining computing nodes are added to the PFS group.
[0045] The application programs in the HPC system are mostly parallel computing programs. A large parallel computing program usually evenly divides the computing tasks and assigns them to multiple processes when running. These processes usually run in parallel on different computing nodes. When a checkpoint is taken, each process needs to save its own computing state. A checkpoint of the entire application is actually multiple independent checkpoints of the processes. The parallel application program is usually load balanced, and the checkpoint data volume of the processes remains consistent. The embodiment is specifically for the case where the checkpoint data volume of each process is the same.
[0046] The storages used by the HeteroCheck for checkpoint caching include SBB and PFS. SBB and PFS are shared storages between multiple applications and multiple nodes in the HPC system. Their capacity is usually sufficient, but the bandwidth they can provide is related to the concurrency degree of the application when running. After the computing nodes are divided into the SBB group and the PFS group in the HeteroCheck framework, to minimize the time for caching the checkpoints, it is necessary to ensure that the SBB group and the PFS group complete the writing of the checkpoints at the same time, otherwise the "tail latency" may occur, that is, one group has completed the writing while the other group has not. At this time, the storage corresponding to the group that has completed the writing is idle and not utilized, and the caching time of the checkpoint is determined by the group that has not completed the writing. On the basis of dividing the computing nodes into the SBB group and the PFS group, the number of nodes that concurrently write to the PFS group and the number of nodes that concurrently write to the SBB group are further controlled in the embodiment to minimize the time difference between the completion of the checkpoint writing of the SBB group and the PFS group.
[0047] In this embodiment, the optimal computing node partitioning manner of the SBB group and the PFS group is determined so that the SBB group and the PFS group complete checkpoint writing with the minimum time difference, which includes: The following computing node partitioning model is constructed: F( )=∣ ∣ (1) wherein, is the number of nodes concurrently writing to the PFS, is the number of nodes concurrently writing to the SBB, is the total amount of data written to the SBB, is the total amount of data written to the PFS, , are the bandwidths of the SBB and the PFS, respectively, when the concurrency is , is the total writing time of the SBB group, is the total writing time of the PFS group, and F( ) is a function of and , and represents the absolute difference between the writing times of the SBB group and the PFS group; By finding the solution that makes F( ) minimum, the optimal computing node partitioning manner that makes the SBB group and the PFS group complete checkpoint writing with the minimum time difference is obtained. That is, the optimal solution of formula (1) is the solution that makes F( ) minimum. By minimizing the difference F( ), when F( ) is minimum, the node allocation scheme that makes the writing times of the SBB group and the PFS group closest can be found, so that the SBB group and the PFS group can complete checkpoint writing in the closest time, thereby the time for caching the checkpoint can be maximally shortened, and the "tail latency" effect can be avoided.
[0048] As a preferred embodiment, the following constraint conditions can be further provided: (2) (3) (4) (5) (6) (7) wherein, D is the total amount of checkpoint data of the job, NThe number of nodes used for job execution. It is about The function, the functional relationship is used express, It is about The function, the functional relationship is used express.
[0049] In searching for F( In the process of finding the minimum solution, the bandwidth of SBB and PFS under different concurrency levels is obtained to form a concurrency-SBB bandwidth table and a concurrency-PFS bandwidth table. These tables are then traversed to find all possible bandwidth values that satisfy the constraints and the corresponding F(t) is calculated. The multiple F( values) obtained from the search will be used to find the F( values). In the value, F( The value corresponding to the minimum value and The values are used as the number of computing nodes for the SBB group and the PFS group, respectively.
[0050] This embodiment uses the above method to determine the optimal number of nodes for the SBB group and the PFS group. The computing node with the smallest number is added to the SBB group in sequence by sequential selection until the number of nodes required by the SBB group is met. The remaining computing nodes are added to the PFS group, so that the two groups of nodes can complete the checkpoint write as simultaneously as possible, thereby minimizing the tail delay effect.
[0051] Specifically, and This is a function of concurrency. HPC systems can use I / O benchmark programs such as IOR to measure the bandwidth of SBB and PFS at different concurrency levels, thus creating concurrency-SBB bandwidth tables and concurrency-PFS bandwidth tables respectively. This embodiment solves for F( In the process of finding the minimum value, firstly, the bandwidths of SBB and PFS under different concurrency levels are obtained, forming a concurrency-SBB bandwidth table and a concurrency-PFS bandwidth table that correspond to the concurrency level and SBB bandwidth. Then, under the condition of satisfying the constraints, the two tables are traversed to determine all possible bandwidth values. Substitute all possible bandwidth values into equation (1) to calculate the corresponding F( ) values, resulting in multiple F( ) value, in multiple F( Find the minimum value among the values, and at this time the corresponding and The value of SBB group and PFS group is the number of computing nodes respectively. Further, if there are multiple equal minimum values, the first found minimum value can be taken as the number of nodes for concurrent writing of SBB group and PFS group.
[0052] In this embodiment, in the checkpoint caching stage, the checkpoint writing to SBB includes: The client first checks whether there is an unfinished checkpoint flushing operation, and if there is, waits for the completion of the checkpoint flushing operation, otherwise starts the current checkpoint caching process. After starting the current checkpoint caching process, the client registers each state that needs to be cached by the process as an RDMA memory region; The client sends a checkpoint caching RPC request to the corresponding server for each state to be cached, and the request carries the starting address and length information of the current state data in the process memory; The client is blocked so that the current state to be cached is not modified until a positive response to the RPC request is received from the server; After receiving a positive response from the server about all states, the client determines that the SBB checkpoint caching is complete, unblocks the client, and continues to execute the computing task.
[0053] Further, it also includes the server sending a positive response to the RPC request sent by the client, and the steps include: ① The server pre-creates a fixed-size buffer and registers it as an RDMA memory region for each received client RPC request about checkpoint caching.
[0054] If the actual length of each client state data is dynamically allocated, when processing a state with a significantly larger length, a super-large capacity buffer needs to be applied, which is easy to exhaust the available memory of the SBB server node, resulting in the risk of out-of-memory (OOM). By adopting the fixed size strategy, the resource management problem when multiple clients concurrently request can be solved.
[0055] ② The server initializes the RDMA operation counter cnt, which represents the number of operations required to complete the caching of one state of the current client.
[0056] Specifically, the formula for calculating the number of operations required for one state caching of the client can be expressed as: , wherein state_length is the state data length, and buf_size is the server fixed buffer size.
[0057] In a specific application embodiment, the RDMA operation counter cnt can be initialized to the number of operations required by the client to cache one state buffer, calculated in the above manner, and then decremented by 1 each time a cache operation is completed. When the RDMA operation counter cnt is 0, it is determined that the current state data cache is complete.
[0058] ③The server uses RDMA to bypass the operating system kernel to read the state from the memory of the corresponding client process into the server RDMA memory region (a fixed-size buffer) according to the start address and length of the state. RDMA can improve network throughput and data transmission efficiency.
[0059] ④The server writes the state in the buffer to the local storage device (such as an NVMe SSD) and updates the value of the RDMA operation counter cnt.
[0060] ⑤Determine whether the value of the RDMA operation counter reaches a preset value. If not, re-execute the server reading the state from the RDMA memory region according to the start address and length of the current state from the memory of the corresponding client process using RDMA to bypass the operating system kernel, and writing the state in the RDMA memory region to the local storage device, and updating the value of the RDMA operation counter cnt. If yes, it is determined that all state data of the current RPC request has been read and written to the local storage device, and the server positively replies to the corresponding client RPC request.
[0061] Specifically, the server checks whether the value of the RDMA operation counter cnt reaches a preset value, for example, whether it is reduced to 0. If yes, steps ③ and ④ are repeatedly executed. Otherwise, it is determined that all data has been read and written to the local NVMe SSD, and the server positively replies to the corresponding client RPC request. Further, after receiving the positive reply about all states, the client can consider that the SBB checkpoint cache is complete, unblock, and continue to execute the computing task.
[0062] In the checkpoint cache phase of the present embodiment, the checkpoint of the PFS group is cached, including: The PFS group computing node directly writes the checkpoint to the PFS through the I / O interface of the parallel file system.
[0063] In the checkpoint flush phase of the present embodiment, the SBB group uses the client-server mode to flush the data cached in the SBB to the PFS, including: Each process on each computing node of the SBB group sends an RPC request for checkpoint flush to the corresponding server; The client returns immediately after sending the RPC request for checkpoint flush to the server, and continues to execute the computing task; After receiving the checkpoint flush request, the server creates a buffer and reads the checkpoint cache corresponding to the client in the local NVMe SSD into the buffer, and then uses the file system interface provided by the PFS to flush the data in the buffer to the PFS; After flushing all the checkpoint caches of a client to the PFS, the server positively replies to the RPC request of the corresponding client. The client checks whether the server positively replies to the RPC request to determine whether the last checkpoint flush is completed when the client performs the next checkpoint cache.
[0064] Further preferably, in the checkpoint flush stage, the computing nodes of the PFS group and the SBB group keep synchronization when performing the computing task, that is, when the computing nodes of the SBB group send the RPC request for checkpoint flush to the server, the computing nodes of the PFS group block the computing, and when the computing nodes of the SBB group continue to perform the computing task, the computing nodes of the PFS group unblock and continue to perform the computing. Through the synchronization mechanism, the consistency of the computing states of the processes of the application program can be ensured.
[0065] Further preferably, in the checkpoint flush stage, the connection established between each client and the SBB server in the checkpoint cache stage is also maintained, so as to ensure that the client uses the service from the same server in the two stages of checkpoint cache and flush, and the server can find the correct cache in the flush stage.
[0066] In the present embodiment, in the restart stage, the checkpoints of all the application processes are read from the PFS to the corresponding regions in the process address space where the critical states are located, so as to recover the computing of the application processes before the failure. Specifically, the checkpoint cache of the "SBB group" is flushed to the PFS, and the checkpoint cache of the "PFS group" does not need to be flushed because it is directly stored in the PFS. Only when the cache in the SBB is flushed, can the checkpoints of all the application processes be found in the PFS. When the application processes encounter system failure and need to be restarted, the HeteroCheck reads the checkpoints of all the application processes from the PFS to the corresponding regions in the process address space where the critical states are located, so as to recover the computing of the application processes before the failure.
[0067] In summary, the application fuses the two heterogeneous storages of PFS and SBB in the HPC system as the cache of the checkpoint, makes full use of the two heterogeneous storage resources of SBB and PFS in the HPC system to cache the checkpoint, disperses the I / O pressure of the checkpoint, avoids the performance bottleneck of the single storage, shortens the time spent for caching the checkpoint, optimizes the flushing of the SBB cache, adopts the "client-server" mode to interact with the remote heterogeneous storage SBB instead of accessing the SBB through the distributed file system, when the checkpoint cached in the SBB needs to be flushed to the PFS, the computing node sends a request as a client, the SBB receives the request as a server, and directly flushes the data to the PFS, reduces the participation of the computing node in the checkpoint flushing process, does not need the computing node to transfer the data, and significantly improves the checkpoint flushing efficiency.
[0068] The embodiment further provides an electronic device, including a processor and a memory, the memory is used to store a computer program, and the processor is used to execute the computer program to perform the method.
[0069] It can be understood that the method can be executed by a single device, such as a computer or a server, and can also be applied to a distributed scenario to be completed by multiple devices in cooperation. In the distributed scenario, one of the multiple devices can only execute one or more steps in the method, and the multiple devices interact to complete the method. The processor can be implemented in the form of a general-purpose CPU, a microprocessor, an application-specific integrated circuit, or one or more integrated circuits, and is used to execute related programs to implement the method. The memory can be implemented in the form of a read-only memory (ROM), a random access memory (RAM), a static storage device, and a dynamic storage device. The memory can store an operating system and other application programs, and when the method is implemented by software or firmware, the related program codes are stored in the memory and executed by the processor.
[0070] The embodiment further provides a computer readable storage medium storing a computer program, and the computer program is executed by a processor to implement the method.
[0071] Those skilled in the art will appreciate that the embodiments of the present application described above can be provided as a method, system, or computer program product. Accordingly, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer readable storage media (including, but not limited to, disk memory, CD-ROMs, optical storage devices, etc.) embodying computer readable program code. The present application is described in terms of flowcharts and / or block diagrams illustrating the operations according to embodiments of the present application. It will be understood that each block of the flowchart and / or block diagrams, and combinations of blocks in the flowchart and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processing system, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more flowcharts and / or block diagrams. Figure 1 one or more flowcharts and / or block diagrams. Figure 1 one or more flowcharts and / or block diagrams. Figure 1 one or more flowcharts and / or block diagrams. Figure 1 one or more flowcharts and / or block diagrams. Figure 1 one or more flowcharts and / or block diagrams.
[0072] The above merely preferred embodiments of the present application and not intended to limit the present application in any form. Although the present application has been disclosed as above with the preferred embodiments, the present application is not intended to be limited thereto. Therefore, any simple modification, equivalent change and modification of the above embodiments, which do not depart from the technical solution of the present application, according to the technical essence of the present application, shall fall within the scope of protection of the technical solution of the present application.
Claims
1. An asynchronous checkpoint cache control method for a high-performance computing system, characterized in that the steps include... include: The compute nodes in the HPC system are divided into SBB group and PFS group. The SBB group uses SBB as checkpoint cache and the PFS group uses PFS as checkpoint cache. The optimal compute node partitioning method for the SBB group and PFS group is determined so that the SBB group and PFS group complete the checkpoint write with the minimum time difference. A server is started on each SBB node, and a client is embedded in each process on each compute node of the SBB group. Each client communicates with an SBB server. During the checkpoint caching phase, checkpoints for the SBB group and PFS group are written to the SBB and PFS respectively using different modes. For the SBB group, the client-server mode is used to write checkpoints to the SBB, and for the PFS group, the file I / O mode is used to write checkpoints to the PFS. During the checkpoint refresh phase, checkpoints in the cache are refreshed to PFS. Checkpoints already cached in PFS remain unchanged, while checkpoints cached in SBB are refreshed to PFS using a client-server model.
2. The asynchronous checkpoint cache control method for a high-performance computing system according to claim 1, characterized in that, The step of determining the optimal computation node partitioning method for the SBB group and the PFS group to ensure that the SBB group and the PFS group complete the checkpoint write with the minimum time difference includes: Construct the following computation node partitioning model: F( )=∣ ∣ in, The number of nodes that concurrently write to PFS. The number of nodes that concurrently write to the SBB. The total amount of data written to the SBB. The total amount of data written to PFS. , They are respectively at a concurrency level of The bandwidth of SBB at a concurrency level of The bandwidth of PFS at that time This indicates the total write time for group SBB. F( represents the total write time of the PFS group) (Regarding) and The function represents the absolute difference in write time between the SBB group and the PFS group; By finding F( The solution that yields the minimum value is obtained, which leads to the optimal computation node partitioning method that enables the SBB group and PFS group to complete the checkpoint writing with the minimum time difference.
3. The asynchronous checkpoint cache control method for a high-performance computing system according to claim 2, characterized in that, This also includes setting the following constraints: in, D This represents the total number of checkpoint data points for the operation. N The number of nodes used for job execution. It is about The function, the functional relationship is... express, It is about The function, the functional relationship is... express; In searching for F( In the process of finding the minimum solution, the bandwidth of SBB and PFS under different concurrency levels is obtained to form a concurrency-SBB bandwidth table and a concurrency-PFS bandwidth table. These tables are then traversed to find all possible bandwidth values that satisfy the constraints and the corresponding F(t) is calculated. The multiple F( values) obtained from the search will be used to find the F( values). In the value, F( The value corresponding to the minimum value and The values are used as the number of computing nodes for the SBB group and the PFS group, respectively.
4. The asynchronous checkpoint cache control method for a high-performance computing system according to claim 1, characterized in that, During the checkpoint caching phase, writing checkpoints to the SBB using a client-server model includes: The client first checks if there are any incomplete checkpoint refresh operations. If so, it waits for the checkpoint refresh operation to be completed; otherwise, it starts the current checkpoint caching process. After starting the current checkpoint caching process, the client registers each state that its process needs to cache as an RDMA memory region; For each state to be cached, the client sends a checkpoint cache RPC request to the corresponding server, carrying the starting address and length information of the current state data in the process memory. The client is blocked to prevent the current cached state from being modified until a positive response to the RPC request is received from the server. After receiving a positive response regarding all states, the client determines that the SBB checkpoint caching is complete, unblocks the client, and continues executing the computation task.
5. The asynchronous checkpoint cache control method for a high-performance computing system according to claim 4, characterized in that, This also includes the server's positive response to RPC requests sent by the client, including: In response to RPC requests from clients regarding checkpoint caching, the server pre-creates a fixed-size buffer and registers it as an RDMA memory region. The server initializes the RDMA operation counter cnt to represent the number of operations required to complete one state buffer for the current client; The server uses RDMA to bypass the operating system kernel, reads the state from the memory of the corresponding client process according to the starting address and length of the current state, writes the state in the RDMA memory region into the local storage device, and updates the value of the RDMA operation counter cnt. Determine whether the value of the RDMA operation counter has reached the preset value. If not, otherwise re-execute the process whereby the server uses RDMA to bypass the operating system kernel, reads the state from the memory of the corresponding client process according to the starting address and length of the current state in the RDMA memory region, writes the state in the RDMA memory region into the local storage device, and updates the value of the RDMA operation counter cnt. If yes, it is determined that all state data of the current RPC request has been read and written to the local storage device, and the server gives an affirmative response to the corresponding client RPC request.
6. The asynchronous checkpoint cache control method for a high-performance computing system according to any one of claims 1 to 5, characterized in that, During the checkpoint refresh phase, the process of using a client-server model to refresh the data cached in the SBB to the PFS includes: The client corresponding to each process on each compute node of the SBB group sends an RPC request for checkpoint refresh to the corresponding server; After the client sends an RPC request for checkpoint refresh to the server and returns, it continues to execute the computation task. After receiving a checkpoint refresh request, the server creates a buffer and reads the checkpoint cache corresponding to the client from the local storage device into the buffer. Then, it uses the file system interface provided by PFS to refresh the data in the buffer to PFS. After the server flushes all checkpoint caches of a client to PFS, it responds positively to the corresponding client RPC requests. When the client caches the next checkpoint, it checks whether the server has responded positively to the RPC request to determine whether the most recent checkpoint refresh has been completed.
7. The asynchronous checkpoint cache control method for a high-performance computing system according to claim 6, characterized in that, The checkpoint refresh phase also includes maintaining the connection established between each client and the SBB server during the checkpoint caching phase, so that the client can use services from the same server in both the checkpoint caching and refresh phases. During the checkpoint refresh phase, the compute nodes of the PFS group and the SBB group also maintain synchronization while performing computation tasks. That is, when the SBB group compute node sends a checkpoint refresh RPC request to the server, the PFS group compute node blocks computation. When the SBB group compute node continues to perform computation tasks, the PFS group compute node unblocks to continue computation.
8. The asynchronous checkpoint cache control method for a high-performance computing system according to any one of claims 1 to 5, characterized in that, It also includes reading checkpoints from PFS for all application processes during the restart phase, down to the corresponding region in the address space of the process in the specified state, to restore the application processes' computations before the failure.
9. An electronic device comprising a processor and a memory, the memory being used to store a computer program, characterized in that, The processor is used to execute the computer program to perform the method as described in any one of claims 1 to 8.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Performance calculation method and system for high-performance computer workflow scheduling and medium
CN115587014A
Layered caching method and system and related components
CN116048425A
Accelerating shared file checkpoint with local burst buffers
US20190332318A1