This application discloses an asynchronous checkpoint caching control method, device, and medium for high-
performance computing systems. The method includes the following steps: dividing the computing nodes in the
system into SBB groups and PFS groups, and ensuring that the SBB and PFS groups complete checkpoint writing with minimal
time difference; starting a
server on each SBB node, embedding a
client in each process of each computing node in the SBB group, and having each
client communicate with an SBB
server; during the checkpoint caching phase, using different
modes to write checkpoints from the SBB and PFS groups to the corresponding SBB and PFS groups, respectively, where the SBB group uses a
client-
server mode and the PFS group uses a file I / O mode; during the checkpoint refresh phase, checkpoints in the SBB are refreshed using the client-server mode. This application can effectively alleviate performance bottlenecks during large-scale checkpoint applications in HPC systems, improve checkpoint efficiency, and reduce the
impact on computing tasks.