Process system and method for resource-efficient, gradual, asynchronous checkpointing and recovery of data

The process system addresses the inefficiencies of existing checkpointing methods by enabling asynchronous, self-sufficient checkpointing and recovery in distributed applications, resulting in reduced recovery time and resource utilization.

WO2025131313A1PCT designated stage expired Publication Date: 2025-06-26HUAWEI TECH CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2023/087706
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing checkpointing methodologies for distributed applications require all processes to participate in the recovery process, leading to inefficient utilization of resources and increased recovery time.

Method used

A process system that enables asynchronous, gradual, and self-sufficient checkpointing and recovery, where each process can independently reload its latest checkpoint data and replay incoming messages without requiring the involvement of other processes.

Benefits of technology

This approach reduces recovery time and resource usage by allowing failed processes to recover autonomously, minimizes message logging overhead, and enables the temporary release of resources allocated to unaffected processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2023087706_26062025_PF_FP_ABST
    Figure EP2023087706_26062025_PF_FP_ABST
Patent Text Reader

Abstract

A process system (100) performs a checkpoint and recovery using a first set of computing resources (102) and a second set of computing resources (106). The first set of computing resources (102) includes a first controller (104) that is configured to execute a first process, and the second set of computing resources (106) includes a second controller (108) that is configured to execute a second process. The first controller (104) is configured to (i) checkpoint the first process asynchronously with the second process, (ii) store incoming messages from the second process, (iii) determine failure in the first process, (iv) reload the latest checkpoint data, (v) redo local computations, and replay the stored incoming messages. The second controller (108) is configured to (a) perform local computations and / or send messages to the first process, (b) determine the failure, and (c) suspend the second process, freeing the second set of computing resources (106).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] PROCESS SYSTEM AND METHOD FOR RESOURCE-EFFICIENT, GRADUAL, ASYNCHRONOUS CHECKPOINTING AND RECOVERY OF DATA

[0002] TECHNICAL FIELD

[0003] The disclosure generally relates to checkpointing schemes, and more particularly, the disclosure relates to a process system configured to checkpoint asynchronously and recover local data. The disclosure also relates to a method for a process system configured to checkpoint asynchronously and recover local data.

[0004] BACKGROUND

[0005] Computer applications have become an integral part of various fields, ranging from High-Performance Computing, HPC to Artificial Intelligence, Al. With the increasing need for computing power, these applications are now being distributed across multiple nodes or processors to efficiently handle the growing workloads. This technique is called single-program multipledata, SPMD parallelism, where the computations are divided into identical processes that operate on individual data partitions. To handle large datasets and complex computations efficiently, large-scale computing clusters are recommended. These clusters consist of interconnected nodes or servers and are used for scientific simulations, data analytics, machine learning, and other purposes.

[0006] However, many applications that run on these compute clusters can take a considerable amount of time, ranging from hours, days, or even weeks, and are prone to hardware or software failures. Failure is an unexpected crash in a single process or in a group of processes in a cluster. Given the complexity of the environment, it is essential to ensure that work progress is resumed with minimal penalty in the event of such failures. Hardware failures, network disruptions, and software glitches are common in these large-scale clusters as they involve multiple components. The more significant the number of components, the higher the likelihood of failure, and the meantime between failures, MTBF, is reduced. Therefore, it is critical to have robust strategies to handle these failures and minimize data loss.

[0007] One of the most vital strategies in this context is checkpointing. Checkpointing involves periodically saving the current state of an ongoing application or job to non-volatile memory / storage for later recovery in case of failure. Due to the high failure rate in high-performance computing, HPC, and distributed computing environments, successful checkpointing schemes have become crucial. In the event of a failure, the saved checkpoint enables the computation to resume from where it left off, rather than starting from scratch. This is a fundamental element in ensuring fault tolerance, resource efficiency, job resilience, cost savings, and the uninterrupted progress of scientific work within these complex computing environments.

[0008] The main objective behind implementing checkpointing is to ensure fault tolerance. In the event of a node or component failure within a computing system, the work completed since the last checkpoint is usually at risk of being lost. The checkpointing aids in preserving the state of the computation at regular intervals, thereby minimizing the penalty incurred in case of failure. In the event of a failure, the capability to resume the computation from the last checkpoint is enabled, thereby saving the need to restart from the beginning. The message passing interface, MPI standard does not currently define fault tolerance requirements. Consequently, most MPI implementations do not support this feature, so by default, the entire MPI application is aborted upon a single process failure.

[0009] Furthermore, checkpointing enhances resource efficiency, as it helps reduce resource wastage by allowing the restart of failed jobs without starting from scratch. It also contributes to job resilience by providing a safety net, reducing the impact of job failures on productivity. Moreover, it results in cost savings by reducing the need for redundant computations and minimizing downtime caused by failures.

[0010] In scientific computing, experiments and simulations often involve long-running jobs. Checkpointing plays a crucial role in safeguarding the continuity of scientific advancement, preventing the loss of progress attributable to hardware or software glitches. This importance is particularly pronounced in the advancement of research within domains like climate modeling, drug discovery, and astrophysics.

[0011] Users of large-scale clusters appreciate the reliability and predictability that checkpointing provides, contributing to the preservation of a streamlined and effective workflow. This reliability is of paramount significance in academic, industrial, and research environments.

[0012] Several checkpointing methodologies are available for distributed message passing interface, MPI applications. The checkpointing methodologies typically require significant resources. One of the existing checkpointing methodologies provides a local rollback approach. The local rollback approach enables the recovery of only the processes that have encountered failures from the most recent checkpoint, ensuring consistency in the execution's progress through a two-tier message logging process. However, a notable drawback of this method is its requirement for the participation of all processes in the recovery process, which is necessary to aid in the restoration of the failed processes. Consequently, after a failure, surviving processes cannot be completely suspended, and have to wait for the failing processes to restart and catch-up. During this time, any resources they were using cannot be freed and reallocated to other jobs in the system. There is no possibility of a self-sufficient recovery of failed processes while other active processes are suspended to free their resources until the recovery is complete. Other checkpointing methods utilize message logging to allow for the recovery of a failed process locally. However, all the methods suggested so far require all distributed processes to participate in the recovery process, which hinders the efficient utilization of idle resources during the recovery process.

[0013] Therefore, there arises a need to address the aforementioned technical problem / drawbacks of selecting an appropriate backup policy for an application.

[0014] SUMMARY

[0015] It is an object of the disclosure to provide a process system configured to checkpoint asynchronously and recover local data and a method for a process system configured to checkpoint asynchronously and recover local data while avoiding one or more disadvantages of prior art approaches.

[0016] This object is achieved by the features of the independent claims. Further, implementation forms are apparent from the dependent claims, the description, and the figures.

[0017] According to a first aspect, there is a process system that includes a first set of computing resources including a first controller configured to execute a first process and a second set of computing resources including a second controller configured to execute a second process. The first controller is configured to perform a checkpoint for the first process asynchronously with the second process. The first controller is configured to store incoming messages from the second process to the first process. The first controller is configured to determine that a failure has occurred in the first process and in response thereto, the first controller is configured to reload latest checkpoint data for the first process, redo local computations and replay the stored incoming messages. The second controller is configured to execute the second process including performing local computations and / or sending messages to the first process. The second controller is configured to determine that the failure has occurred in the first process and in response thereto, the second controller is configured to suspend the second process thereby freeing up the second set of computing resources.

[0018] The process system offers a resource-efficient, gradual, and asynchronous checkpointing method for distributed applications. The asynchronous checkpointing method of the process system effectively alleviates the input / output I / O, network, and storage burdens associated with synchronous checkpointing through a gradual method that leverages message passing interface, MPI and asynchronous I / O capabilities in order to orchestrate checkpoints among individual processes. Moreover, the asynchronous checkpointing method enables the independent self-sufficient recovery of failed processes in single-program, multiple-data, SPMD distributed applications. A combined approach of checkpointing and incoming MPI message logging ensures comprehensive self-sufficient recovery capabilities throughout the SPMD application's entire lifecycle, where the failed processes can autonomously 'recover by themselves' after a failure without necessitating the involvement of other processes. During recovery, the failed process only needs access to the latest local checkpoint and an incoming message log, as all nonfailed processes are suspended and the corresponding resources may be freed for use by other jobs, thus being self-sufficient. Furthermore, the process system enables the avoidance of logging sent messages, substantially reducing message logging overhead. Most importantly, the process system facilitates the temporary release of resources allocated to unaffected processes until the failed process is back in operation. The process system encompasses enabling a resilient approach for demanding applications like high-performance computing, HPC and artificial intelligence, Al workloads. The process system enables the reduction of recovery time and necessary resources through self-sufficient recovery. The process system minimizes message logging overhead by exclusively recording incoming messages while reproducing sent messages as outcomes of local computations. The process system may be applied in the distributed training of Al and machine learning, ML models, involving the utilization of multiple worker nodes to train the model on distinct data sets, a process known as data parallel training.

[0019] The process system elevates the reliability, fault tolerance, job resilience, resource efficiency, cost savings, scientific progress, user experience, and overall performance of the resource-intensive tasks. Hence, the process system ensures that these applications operate smoothly even in challenging conditions or when encountering unexpected errors. One of the process system's standout features is the complete self-sufficient recovery capability, guaranteeing that failed processes can autonomously recover after a failure without requiring the involvement of other processes. During recovery, each failed process only needs access to the latest local checkpoint and the incoming message log. As a result, the process system not only ensures swift recovery but also significantly reduces the overhead from message logging. The process system allows the avoidance of logging sent messages, thus improving the overall system efficiency. To further assess the process system's efficiency, various checks can be performed, including examining the I / O pattern throughout normal runtime and recovery of the application to see if and when message logs are created or accessed, verifying if only incoming messages are logged by comparing the message log size to the volume of outgoing and incoming communication traffic, and forcing a process to fail to determine whether the remaining processes are suspended or remain active to assist in the bring-up.

[0020] The process system is resilient to failure scenarios. Furthermore, the process system ensures the checkpointing is done asynchronously without disrupting the operations. The comprehensive approach of the process system optimizes resource utilization and fault tolerance, making the process system highly suitable for resource-constrained distributed environments. In summary, the process system combines adaptability, self-sufficient recovery, and efficient message logging to deliver a robust solution for HPC and Al workloads.

[0021] Optionally, the first controller is further configured to store incoming messages from the second process to the first process received after the checkpoint. Optionally, the first controller is further configured to send a signal to the second process indicating that the first process has failed. The second controller is further configured to determine that the failure in the first process has occurred by receiving the signal that a failure has occurred in the first process.

[0022] Optionally, the first controller is further configured to send a second signal to the second process indicating the successful recovery of the first process. The second controller is further configured to receive the second signal indicating the successful recovery of the first process, and in response thereto wake the second process, reclaim required resources, and continue its normal execution.

[0023] Optionally, the first controller is further configured to abstain from sending outgoing messages while redoing local instructions during the recovery stage.

[0024] Optionally, the first controller is further configured to replay all the stored incoming messages since the latest check point, in the order that they were stored in.

[0025] Optionally, the first controller is further configured to determine that a point of failure has been reached while redoing local instructions and in response thereto cause the second controller to rewind the second process back to the latest checkpoint which serves as a global recovery point.

[0026] Optionally, the first and second controllers are configured to perform the checkpoint in a gradual and exclusive manner, such that two processes are not performing the checkpoint at a same time.

[0027] Optionally, the first controller is further configured to perform the checkpoint for the first process at a first time and the second controller is configured to perform the checkpoint for the second process at a second time, the first time is different from the second time.

[0028] Optionally, the first time and the second time do not overlap one another.

[0029] Optionally, the process system further includes an orchestrator, the orchestrator is configured to indicate to a process when to perform the checkpoint.

[0030] Optionally, the orchestrator is configured to indicate to a process when to perform the checkpoint by assigning the first time and the second time.

[0031] Optionally, the orchestrator is further configured to assign the first time and the second time in order of process identifier for the first and second processes.

[0032] Optionally, the orchestrator is configured to assign the first time and the second time based on a number of operations performed by each of the first and second processes.

[0033] Optionally, the orchestrator is configured to assign the first time and the second time based on a number of memory operations performed by each of the first and second processes.

[0034] Optionally, the orchestrator is configured to assign the first time and the second time based on the volume of checkpoint data stored by each of the first and second processes.

[0035] According to a second aspect, there is provided a first arrangement of computing resources including a first controller configured to execute a first process, and a second set of computing resources including a second controller configured to execute a second process. The first controller is configured to perform a checkpoint for the first process. The first controller is configured to store incoming messages to the first process. The first controller is configured to determine that a failure has occurred and in response thereto, the first controller is configured to reload the checkpoint for the first process. The first controller is configured to. The first controller is configured to redo local instructions starting from the checkpoint. The first controller is configured to replay the stored incoming messages.

[0036] The first arrangement of computing resources and the second set of computing resources offer a resource-efficient, gradual, and asynchronous checkpointing method for distributed applications. The asynchronous checkpointing method of the process system effectively alleviates the input / output I / O, network, and storage burdens associated with checkpointing through a gradual method that leverages message passing interface, MPI and asynchronous I / O capabilities in order to orchestrate checkpoints among individual processes. Moreover, the asynchronous checkpointing method enables the independent self- sufficient recovery of failed processes in single-program, multiple-data SPMD distributed applications. A combined approach of checkpointing and incoming MPI message logging ensures comprehensive self-sufficient recovery capabilities throughout the SPMD application's entire lifecycle, where the failed processes can autonomously 'recover by themselves' after a failure without necessitating the involvement of other processes. During recovery, the failed process only needs access to latest local checkpoint and an incoming message log, as all non-failed processes are suspended and the corresponding resources may be freed for use by other jobs, thus being self-sufficient. Furthermore, the first arrangement of computing resources and the second set of computing resources enable the avoidance of logging sent messages, substantially reducing message logging overhead. Most importantly, the first arrangement of computing resources and the second set of computing resources facilitates the temporary release of resources allocated to unaffected processes until the failed process is back in operation. The first arrangement of computing resources and the second set of computing resources enable a resilient approach for demanding applications like high-performance computing, HPC and artificial intelligence, Al workloads. The first arrangement of computing resources and the second set of computing resources enable the reduction of recovery time and necessary resources through self-sufficient recovery. The first arrangement of computing resources and the second set of computing resources minimize message logging overhead by exclusively recording incoming messages while reproducing sent messages as outcomes of local computations. The first arrangement of computing resources and the second set of computing resources may be applied in the distributed training of Al and machine learning, ML models, involving the utilization of multiple worker nodes to train the model on distinct data sets, a process known as data parallel training.

[0037] The first arrangement of computing resources and a second set of computing resources elevate the reliability, fault tolerance, job resilience, resource efficiency, cost savings, scientific progress, user experience, and overall performance of the resourceintensive tasks. Hence, the process system ensures that these applications operate smoothly even in challenging conditions or when encountering unexpected errors.

[0038] According to a third aspect, there is provided a second arrangement of computing resources comprising a second controller configured to execute a second process. The second controller is configured to determine that a failure has occurred for a first process and in response thereto the second controller is configured to suspend the second process thereby freeing up the second set of computing resources while the first process recovers.

[0039] The second arrangement of computing resources ensures comprehensive self-sufficient recovery capabilities throughout the SPMD application's entire lifecycle. This means that failed processes can autonomously 'recover by themselves' after a failure without necessitating the involvement of other processes. During recovery, each failed process only needs access to the latest local checkpoint and the incoming message log. As a result, the second arrangement of computing resources not only ensures swift recovery but also significantly reduces the overhead from message logging. The second arrangement of computing resources allows the avoidance of logging sent messages, thus improving the overall system efficiency. Most importantly, the second arrangement of computing resources facilitates the temporary release of resources allocated to unaffected processes until the failed process is back in operation. The key benefits of the second arrangement of computing resources encompass enabling a resilient approach for demanding applications like HPC and Al workloads, the reduction of recovery time and necessary resources through self-sufficient recovery, and the minimization of message logging overhead by exclusively recording incoming messages while reproducing sent messages as outcomes of local computations.

[0040] The second arrangement of computing resources elevates the reliability, fault tolerance, job resilience, resource efficiency, cost savings, scientific progress, user experience, and overall performance of these resource-intensive tasks. Hence, the second arrangement of computing resources ensures that these applications operate smoothly even in challenging conditions or when encountering unexpected errors. One of the standout features of the second arrangement of computing resources is the complete self-sufficient recovery capability, guaranteeing that failed processes can autonomously recover after a failure without requiring the involvement of other processes. During recovery, each failed process only needs access to the latest local checkpoint and the incoming message log. As a result, the second arrangement of computing resources not only ensures swift recovery but also significantly reduces the overhead from message logging. The second arrangement of computing resources allows the avoidance of logging sent messages, thus improving the overall system efficiency. To further assess the system's efficiency, various checks can be performed, including examining the I / O pattern throughout normal runtime and recovery of the application to see if and when message logs are created or accessed, verifying if only incoming messages are logged by comparing the message log size to the volume of outgoing and incoming communication traffic and forcing a process to fail to determine whether the remaining processes are suspended or remain active to assist in the bring-up.

[0041] The operations of the second arrangement of computing resources enhance resilience to failure scenarios. The comprehensive approach of the second arrangement of computing resources optimizes resource utilization and fault tolerance, making the second arrangement of computing resources highly suitable for resource-constrained distributed environments. In summary, the second arrangement of computing resources combines adaptability, self-sufficient recovery, and efficient message logging to deliver a robust solution for HPC and Al workloads.

[0042] According to a fourth aspect, there is provided a process system, including a first set of computing resources including a first controller configured to execute a first process and a second set of computing resources including a second controller configured to execute a second process. The first controller is configured to perform a checkpoint for the first process. The first controller is configured to store incoming messages from the second process to the first process. The first controller is configured to determine that a failure has occurred in the first process and in response thereto, the first controller is configured to reload the latest checkpoint data for the first process. The first controller is configured to redo local computations, and replay the stored incoming messages, in order to bring-up the first process to the up-to-date (latest) state of the second process. The second controller is configured to execute the second process including performing local computations and / or sending messages to the first process. The second controller is configured to determine that the failure has occurred in the first process and in response thereto, the second controller is configured to suspend the second process thereby freeing up the second set of computing resources.

[0043] Optionally, the first controller is further configured to store incoming messages from the second process to the first process received after the checkpoint.

[0044] The process system offers a resource-efficient, gradual, and asynchronous checkpointing method for distributed applications. The asynchronous checkpointing method of the process system effectively alleviates the input / output I / O, network, and storage burdens associated with synchronous checkpointing through a gradual method that leverages message passing interface, MPI and asynchronous I / O capabilities in order to orchestrate checkpoints among individual processes. Moreover, the asynchronous checkpointing method enables the independent self-sufficient recovery of failed processes in single-program, multiple-data, SPMD distributed applications. A combined approach of checkpointing and incoming MPI message logging ensures comprehensive self-sufficient recovery capabilities throughout the SPMD application's entire lifecycle, where the failed processes can autonomously 'recover by themselves' after a failure without necessitating the involvement of other processes. During recovery, the failed process only needs access to latest local checkpoint and an incoming message log, as all non-failed processes are suspended and the corresponding resources may be freed for use by other jobs, thus being self-sufficient. Furthermore, the process system enables the avoidance of logging sent messages, substantially reducing message logging overhead. Most importantly, the process system facilitates the temporary release of resources allocated to unaffected processes until the failed process is back in operation. The process system encompasses enabling a resilient approach for demanding applications like high-performance computing, HPC and artificial intelligence, Al workloads. The process system enables the reduction of recovery time and necessary resources through self-sufficient recovery. The process system minimizes message logging overhead by exclusively recording incoming messages while reproducing sent messages as outcomes of local computations. The process system may be applied in the distributed training of Al and machine learning, ML models, involving the utilization of multiple worker nodes to train the model on distinct data sets, a process known as data parallel training.

[0045] The process system elevates the reliability, fault tolerance, job resilience, resource efficiency, cost savings, scientific progress, user experience, and overall performance of the resource-intensive tasks. Hence, the process system ensures that these applications operate smoothly even in challenging conditions or when encountering unexpected errors. One of the process system's standout features is the complete self-sufficient recovery capability, guaranteeing that failed processes can autonomously recover after a failure without requiring the involvement of other processes. During recovery, each failed process only needs access to the latest local checkpoint and the incoming message log. As a result, the process system not only ensures swift recovery but also significantly reduces the overhead from message logging. The process system allows the avoidance of logging sent messages, thus improving the overall system efficiency. To further assess the process system's efficiency, various checks can be performed, including examining the I / O pattern throughout normal runtime and recovery of the application to see if and when message logs are created or accessed, verifying if only incoming messages are logged by comparing the message log size to the volume of outgoing and incoming communication traffic and forcing a process to fail to determine whether the remaining processes are suspended or remain active to assist in the bring-up.

[0046] The process system is resilient to failure scenarios. Further, the process system ensures the checkpointing is done asynchronously without disrupting normal execution. The comprehensive approach of the process system optimizes resource utilization and fault tolerance, making the process system highly suitable for resource-constrained distributed environments. In summary, the process system combines adaptability, self-sufficient recovery, and efficient message logging to deliver a robust solution for HPC and Al workloads.

[0047] Optionally, the first controller is further configured to send a signal to the second process indicating that the first process has failed. The second controller is further configured to determine that the failure in the first process has occurred by receiving the signal that a failure has occurred in the first process.

[0048] Optionally, the first controller is further configured to send a second signal to the second process indicating the successful recovery of the first process. The second controller is further configured to receive the second signal indicating the successful recovery of the first process, and in response thereto wake the second process, reclaim required resources, and continue its normal execution.

[0049] Optionally, the first controller is further configured to abstain from sending outgoing messages while redoing local instructions during the recovery stage.

[0050] Optionally, the first controller is further configured to replay all the stored incoming messages since the latest check point, in the order that they were stored in. Optionally, the first controller is further configured to determine that a point of failure has been reached while redoing local instructions and in response thereto cause the second controller to rewind the second process back to the latest checkpoint which serves as a global recovery point.

[0051] According to a fifth aspect, there is provided a process system, including a first set of computing resources including a first controller configured to execute a first process, and a second set of computing resources including a second controller configured to execute a second process. The first and second controllers are configured to perform checkpoints for the first and second processes in a gradual and exclusive manner, such that two processes are not performing a checkpoint at a same time.

[0052] The process system offers a resource-efficient, gradual, and asynchronous checkpointing method for distributed applications. The asynchronous checkpointing method of the process system effectively alleviates the input / output I / O, network, and storage burdens associated with checkpointing through a gradual method that leverages message passing interface, MPI and asynchronous I / O capabilities in order to orchestrate checkpoints among individual processes. Moreover, the asynchronous checkpointing method enables the independent self-sufficient recovery of failed processes in single-program, multiple-data, SPMD distributed applications. A combined approach of checkpointing and incoming MPI message logging ensures comprehensive self-sufficient recovery capabilities throughout the SPMD application's entire lifecycle, where the failed processes can autonomously 'recover by themselves' after a failure without necessitating the involvement of other processes. During recovery, the failed process only needs access to the latest local checkpoint and an incoming message log, as all nonfailed processes are suspended and the corresponding resources may be freed for use by other jobs, thus being self-sufficient. Furthermore, the process system enables the avoidance of logging sent messages, substantially reducing message logging overhead. Most importantly, the process system facilitates the temporary release of resources allocated to unaffected processes until the failed process is back in operation. The process system encompasses enabling a resilient approach for demanding applications like high-performance computing, HPC and artificial intelligence, Al workloads. The process system enables the reduction of recovery time and necessary resources through self-sufficient recovery. The process system minimizes message logging overhead by exclusively recording incoming messages while reproducing sent messages as outcomes of local computations. The process system may be applied in the distributed training of Al and machine learning, ML models, involving the utilization of multiple worker nodes to train the model on distinct data sets, a process known as data parallel training.

[0053] The process system elevates the reliability, fault tolerance, job resilience, resource efficiency, cost savings, scientific progress, user experience, and overall performance of the resource-intensive tasks. Hence, the process system ensures that these applications operate smoothly even in challenging conditions or when encountering unexpected errors. One of the process system's standout features is the complete self-sufficient recovery capability, guaranteeing that failed processes can autonomously recover after a failure without requiring the involvement of other processes. During recovery, each failed process only needs access to the latest local checkpoint and the incoming message log. As a result, the process system not only ensures swift recovery but also significantly reduces the overhead from message logging. The process system allows the avoidance of logging sent messages, thus improving the overall system efficiency. To further assess the process system's efficiency, various checks can be performed, including examining the I / O pattern throughout normal runtime and recovery of the application to see if and when message logs are created or accessed, verifying if only incoming messages are logged by comparing the message log size to the volume of outgoing and incoming communication traffic, and forcing a process to fail to determine whether the remaining processes are suspended or remain active to assist in the bring-up.

[0054] The process system is resilient to failure scenarios. Furthermore, the process system ensures the checkpointing is done asynchronously without disrupting the operations. The comprehensive approach of the process system optimizes resource utilization and fault tolerance, making the process system highly suitable for resource-constrained distributed environments. In summary, the process system combines adaptability, self-sufficient recovery, and efficient message logging to deliver a robust solution for HPC and Al workloads. Optionally, the first controller is further configured to perform the checkpoint for the first process at a first time and the second controller is configured to perform the checkpoint for the second process at a second time, the first time is different from the second time.

[0055] Optionally, the first time and the second time do not overlap one another.

[0056] Optionally, the process system further includes an orchestrator, the orchestrator is configured to indicate to a process when to perform the checkpoint.

[0057] Optionally, the orchestrator is configured to indicate to a process when to perform the checkpoint by assigning the first time and the second time.

[0058] Optionally, the orchestrator is further configured to assign the first time and the second time in order of process identifier for the first and second processes.

[0059] Optionally, the orchestrator is configured to assign the first time and the second time based on a number of operations performed by each of the first and second processes.

[0060] Optionally, the orchestrator is configured to assign the first time and the second time based on a number of memory operations performed by each of the first and second processes.

[0061] Optionally, the orchestrator is configured to assign the first time and the second time based on the volume of checkpoint data stored by each of the first and second processes.

[0062] Optionally, the process system is part of a High-Performance Computing, HPC, system.

[0063] Optionally, the process system is configured to operate utilizing the Message Passing Interface, MPI.

[0064] According to a sixth aspect, there is provided a method for a process system including a first set of computing resources including a first controller configured to execute a first process and a second set of computing resources comprising a second controller configured to execute a second process. The method includes the first controller performing a checkpoint for the first process asynchronously with the second process. The method includes the first controller storing incoming messages from the second process to the first process. The method includes the first controller determining that a failure has occurred in the first process and in response thereto . The method includes the first controller reloading the latest checkpoint data for the first process. The method includes the first controller redoing local computations, and replay the stored incoming messages. The method includes the second controller executing the second process including performing local computations and / or sending messages to the first process. The method includes the second controller determining that the failure has occurred in the first process and in response thereto. The method includes the second controller suspending the second process thereby freeing up the second set of computing resources. The method includes the first and second controllers performing asynchronous checkpoints for the first and second processes in a gradual and exclusive manner, such that two processes are not performing a checkpoint at a same time.

[0065] The method offers a resource-efficient, gradual, and asynchronous checkpointing method for distributed applications. The asynchronous checkpointing method of the method effectively alleviates the input / output I / O, network, and storage burdens associated with synchronous checkpointing through a gradual method that leverages message passing interface, MPI and asynchronous I / O capabilities in order to orchestrate checkpoints among individual processes. Moreover, the asynchronous checkpointing method enables the independent self-sufficient recovery of failed processes in single-program, multiple-data distributed, SPMD applications. Combined approach of checkpointing and incoming MPI message logging ensures comprehensive self-sufficient recovery capabilities throughout the SPMD application's entire lifecycle, where the failed processes can autonomously 'recover by themselves' after a failure without necessitating the involvement of other processes. During recovery, the failed process only needs access to latest local checkpoint and an incoming message log, as all non-failed processes are suspended and the corresponding resources may be freed for use by other jobs, thus being self-sufficient. Furthermore, the method enables the avoidance of logging sent messages, substantially reducing message logging overhead. Most importantly, the method facilitates the temporary release of resources allocated to unaffected processes until the failed process is back in operation. The method encompasses enabling a resilient approach for demanding applications like high- performance computing, HPC and artificial intelligence, Al workloads. The method enables the reduction of recovery time and necessary resources through self-sufficient recovery. The method minimizes message logging overhead by exclusively recording incoming messages while reproducing sent messages as outcomes of local computations. The method may be applied in the distributed training of Al and machine learning, ML models, involving the utilization of multiple worker nodes to train the model on distinct data sets, a process known as data parallel training.

[0066] The method elevates the reliability, fault tolerance, job resilience, resource efficiency, cost savings, scientific progress, user experience, and overall performance of the resource-intensive tasks. Hence, the method ensures that these applications operate smoothly even in challenging conditions or when encountering unexpected errors. One of the method's standout features is the complete self-sufficient recovery capability, guaranteeing that failed processes can autonomously recover after a failure without requiring the involvement of other processes. During recovery, each failed process only needs access to the latest local checkpoint and the incoming message log. As a result, the method not only ensures swift recovery but also significantly reduces the overhead from message logging. The method allows the avoidance of logging sent messages, thus improving the overall system efficiency. To further assess the method's efficiency, various checks can be performed, including examining the I / O pattern throughout normal runtime and recovery of the application to see if and when message logs are created or accessed, verifying if only incoming messages are logged by comparing the message log size to the volume of outgoing and incoming communication traffic, and forcing a process to fail to determine whether the remaining processes are suspended or remain active to assist in the bring-up.

[0067] The method is resilient to failure scenarios. Furthermore, the method ensures the checkpointing is done asynchronously without disrupting the operations. The comprehensive approach of the method optimizes resource utilization and fault tolerance, making the method highly suitable for resource-constrained distributed environments. In summary, the method combines adaptability, self-sufficient recovery, and efficient message logging to deliver a robust solution for HPC and Al workloads.

[0068] According to a seventh aspect, there is provided a method including a first set of computing resources including a first controller configured to execute a first process and a second set of computing resources comprising a second controller configured to execute a second process. The method includes the first controller performing a checkpoint for the first process. The method includes the first controller storing incoming messages from the second process to the first process. The method includes the first controller determining that a failure has occurred in the first process and in response thereto. The method includes the first controller reloading the latest checkpoint data for the first process. The method includes the first controller redoing local computations, and replaying the stored incoming messages. The method includes the second controller executing the second process including performing local computations and / or sending messages to the first process. The method includes the second controller determining that the failure has occurred in the first process and in response thereto. The method includes the second controller suspending the second process thereby freeing up the second set of computing resources.

[0069] The method offers a resource-efficient, gradual, and asynchronous checkpointing method for distributed applications. The asynchronous checkpointing method of the method effectively alleviates the input / output I / O, network, and storage burdens associated with synchronous checkpointing through a gradual method that leverages message passing interface, MPI and asynchronous I / O capabilities in order to synchronize checkpoints among individual processes. Moreover, the asynchronous checkpointing method enables the independent self-sufficient recovery of failed processes in single-program, multiple-data, SPMD distributed applications. A combined approach of checkpointing and incoming MPI message logging ensures comprehensive self-sufficient recovery capabilities throughout the SPMD application's entire lifecycle, where the failed processes can autonomously 'recover by themselves' after a failure without necessitating the involvement of other processes. During recovery, the failed process only needs access to latest local checkpoint and an incoming message log, as all non-failed processes are suspended and the corresponding resources may be freed for use by other jobs, thus being self-sufficient. Furthermore, the method enables the avoidance of logging sent messages, substantially reducing message logging overhead. Most importantly, the method facilitates the temporary release of resources allocated to unaffected processes until the failed process is back in operation. The method encompasses enabling a resilient approach for demanding applications like high- performance computing, HPC and artificial intelligence, Al workloads. The method enables the reduction of recovery time and necessary resources through self-sufficient recovery. The method minimizes message logging overhead by exclusively recording incoming messages while reproducing sent messages as outcomes of local computations. The method may be applied in the distributed training of Al and machine learning, ML models, involving the utilization of multiple worker nodes to train the model on distinct data sets, a process known as data parallel training.

[0070] The method elevates the reliability, fault tolerance, job resilience, resource efficiency, cost savings, scientific progress, user experience, and overall performance of the resource-intensive tasks. Hence, the method ensures that these applications operate smoothly even in challenging conditions or when encountering unexpected errors. One of the method's standout features is the complete self-sufficient recovery capability, guaranteeing that failed processes can autonomously recover after a failure without requiring the involvement of other processes. During recovery, each failed process only needs access to the latest local checkpoint and the incoming message log. As a result, the method not only ensures swift recovery but also significantly reduces the overhead from message logging. The method allows the avoidance of logging sent messages, thus improving the overall system efficiency. To further assess the method's efficiency, various checks can be performed, including examining the I / O pattern throughout normal runtime and recovery of the application to see if and when message logs are created or accessed, verifying if only incoming messages are logged by comparing the message log size to the volume of outgoing and incoming communication traffic, and forcing a process to fail to determine whether the remaining processes are suspended or remain active to assist in the bring-up.

[0071] The method is resilient to failure scenarios. Furthermore, the method ensures the checkpointing is done asynchronously without disrupting the operations. The comprehensive approach of the method optimizes resource utilization and fault tolerance, making the method highly suitable for resource-constrained distributed environments. In summary, the method combines adaptability, self-sufficient recovery, and efficient message logging to deliver a robust solution for HPC and Al workloads.

[0072] According to an eighth aspect, there is provided a method for a process system including a first set of computing resources including a first controller configured to execute a first process, and a second set of computing resources including a second controller configured to execute a second process. The method includes the first and second controllers performing checkpoints for the first and second processes in a gradual and exclusive manner, such that two processes are not performing a checkpoint at a same time.

[0073] According to a ninth aspect, a computer program product includes program instructions for performing the method when executed by one or more processors in a process system.

[0074] Therefore, in contradistinction to the existing solutions, a process system is configured to provide asynchronous checkpointing in a gradual and exclusive manner such that two processes are not performing a checkpoint at a same time. The process system is configured to provide recovery of data by determining a failure in a process and suspending the other process. These and other aspects of the disclosure will be apparent from the implementation s) described below.

[0075] BRIEF DESCRIPTION OF DRAWINGS

[0076] Implementations of the disclosure will now be described, by way of example only, with reference to the accompanying drawings, in which:

[0077] FIG. 1 is a block diagram of a method including a first set of computing resources and a second set of computing resources for checkpointing and recovery of data using a first set of computing resources and a second set of computing resources in accordance with an implementation of the disclosure;

[0078] FIG. 2 illustrates a block diagram of a first arrangement of computing resources including a first controller configured to execute a first process and a second set of computing resources including a second controller configured to execute a second process in accordance with an implementation of the disclosure;

[0079] FIG. 3 illustrates a block diagram of a second arrangement of computing resources including a first controller configured to execute a first process and a second set of computing resources including a second controller configured to execute a second process in accordance with an implementation of the disclosure;

[0080] FIG. 4 is a block diagram of a process system for recovery of data using a first set of computing resources and a second set of computing resources in accordance with an implementation of the disclosure;

[0081] FIG. 5 is a block diagram of a process system for checkpointing using a first set of computing resources and a second set of computing resources in accordance with an implementation of the disclosure;

[0082] FIG. 6 illustrates an exemplary representation of checkpointing and local recovery of data in accordance with an implementation of the disclosure;

[0083] FIGS. 7A-7C are flow diagrams that illustrate a method for a process system to perform a checkpoint and recovery of data using a first set of computing resources and a second set of computing resources in accordance with an implementation of the disclosure;

[0084] FIGS. 8A-8C are flow diagrams that illustrate a method for a process system including a first set of computing resources including a first controller configured to execute a first process and a second set of computing resources including a second controller configured to execute a second process in accordance with an implementation of the disclosure;

[0085] FIG. 9 is a flow diagram that illustrates a method for a including a first set of computing resources including a first controller configured to execute a first process and a second set of computing resources including a second controller configured to execute a second process in accordance with an implementation of the disclosure; and

[0086] FIG. 10 is an illustration of a computer system (e.g., a first controller, second controller) in which the various architectures and functionalities of the various previous implementations may be implemented. DETAILED DESCRIPTION OF THE DRAWINGS

[0087] Implementations of the disclosure provide a process system configured to checkpoint asynchronously and recover local data and a method for a process system configured to checkpoint asynchronously and recover local data.

[0088] To make solutions of the disclosure more comprehensible for a person skilled in the art, the following implementations of the disclosure are described with reference to the accompanying drawings.

[0089] Terms such as "a first", "a second", "a third", and "a fourth" (if any) in the summary, claims, and foregoing accompanying drawings of the disclosure are used to distinguish between similar objects and are not necessarily used to describe a specific sequence or order. It should be understood that the terms so used are interchangeable under appropriate circumstances, so that the implementations of the disclosure described herein are, for example, capable of being implemented in sequences other than the sequences illustrated or described herein. Furthermore, the terms "include" and "have" and any variations thereof, are intended to cover a non-exclusive inclusion. For example, a process, a method, a system, a product, or a device that includes a series of steps or units, is not necessarily limited to expressly listed steps or units but may include other steps or units that are not expressly listed or that are inherent to such process, method, product, or device.

[0090] FIG. 1 is a block diagram of a process system 100 including a first set of computing resources 102 and a second set of computing resources 106 for checkpointing and recovery of data in accordance with an implementation of the disclosure. The first set of computing resources 102 includes a first controller 104 that is configured to execute a first process 110, and the second set of computing resources 106 includes a second controller 108 that is configured to execute a second process 112. The first controller 104 is configured to (i) perform a checkpoint for the first process 110 asynchronously with the second process 112, (ii) store incoming messages from the second process 112 to the first process 110, (iii) determine that a failure has occurred in the first process 110 and in response thereto, (iv) reload the latest checkpoint data for the first process 110, (v) redo local computations, and replay the stored incoming messages. The second controller 108 is configured to (a) execute the second process 112 including performing local computations and / or sending messages to the first process, 110 and (b) determine that the failure has occurred in the first process 110 and in response thereto, suspend the second process 112 thereby freeing up the second set of computing resources 106.

[0091] The asynchronous also refers to writing checkpoint data in a non-blocking manner, i.e. while enabling an application to continue running while data is written to a disk in the background. In addition, the asynchronous also refers to a (lack of) synchronization between processes, i.e., the first process 100 and the second process 112.

[0092] The process system 100 offers a resource-efficient, gradual, and asynchronous checkpointing method for distributed applications. The asynchronous checkpointing method of the process system 100 effectively alleviates the input / output I / O. network, and storage burdens associated with synchronous checkpointing through a gradual method that leverages message passing interface, MPI and asynchronous I / O capabilities in order to orchestrate checkpoints among individual processes. Moreover, the asynchronous checkpointing method enables the independent self-sufficient recovery of failed processes in single-program, multiple-data, SPMD distributed applications. A combined approach of checkpointing and incoming MPI message logging ensures comprehensive self-sufficient recovery capabilities throughout the SPMD application's entire lifecycle, where the failed processes can autonomously 'recover by themselves' after a failure without necessitating the involvement of other processes. During recovery, the failed process only needs access to latest local checkpoint and an incoming message log, as all non-failed processes are suspended and the corresponding resources may be freed for use by other jobs, thus being self-sufficient. Furthermore, the process system 100 enables the avoidance of logging sent messages, substantially reducing message logging overhead. Most importantly, the process system 100 facilitates the temporary release of resources allocated to unaffected processes until the failed process is back in operation. The process system 100 encompasses enabling a resilient approach for demanding applications like high-performance computing, HPC and artificial intelligence, Al workloads. The process system 100 enables the reduction of recovery time and necessary resources through self-sufficient recovery. The process system 100 minimizes message logging overhead by exclusively recording incoming messages while reproducing sent messages as outcomes of local computations. The process system 100 may be applied in the distributed training of Al and machine learning, ML models, involving the utilization of multiple worker nodes to train the model on distinct data sets, a process known as data parallel training.

[0093] The process system 100 elevates the reliability, fault tolerance, job resilience, resource efficiency, cost savings, scientific progress, user experience, and overall performance of the resource-intensive tasks. Hence, the process system 100 ensures that these applications operate smoothly even in challenging conditions or when encountering unexpected errors. One of the process system's standout features is the complete self-sufficient recovery capability, guaranteeing that failed processes can autonomously recover after a failure without requiring the involvement of other processes. During recovery, each failed process only needs access to the latest local checkpoint and the incoming message log. As a result, the process system 100 not only ensures swift recovery but also significantly reduces the overhead from message logging. The process system 100 allows the avoidance of logging sent messages, thus improving the overall system efficiency. To further assess the process system's efficiency, various checks can be performed, including examining the I / O pattern throughout normal runtime and recovery of the application to see if and when message logs are created or accessed, verifying if only incoming messages are logged by comparing the message log size to the volume of outgoing and incoming communication traffic, and forcing a process to fail to determine whether the remaining processes are suspended or remain active to assist in the bring-up.

[0094] The process system 100 is resilient to failure scenarios. Furthermore, the process system 100 ensures the checkpointing is done asynchronously without disrupting the operations. The comprehensive approach of the process system 100 optimizes resource utilization and fault tolerance, making the process system 100 highly suitable for resource-constrained distributed environments. In summary, the process system 100 combines adaptability, self-sufficient recovery, and efficient message logging to deliver a robust solution for HPC and Al workloads.

[0095] Optionally, the first controller 104 is further configured to store incoming messages from the second process 112 to the first process 110 received after the checkpoint.

[0096] Optionally, the first controller 104 is further configured to send a signal to the second process 112 indicating that the first process 110 has failed. The second controller 108 is further configured to determine that the failure in the first process 110 has occurred by receiving the signal that a failure has occurred in the first process 110.

[0097] Optionally, the first controller 104 is further configured to send a second signal to the second process 112 indicating the successful recovery of the first process 110. The second controller 108 is further configured to receive the second signal indicating the successful recovery of the first process 110, and in response thereto wake the second process 112, reclaim required resources, and continue its normal execution.

[0098] Optionally, the first controller 104 is further configured to abstain from sending outgoing messages while redoing local instructions during the recovery stage.

[0099] Optionally, the first controller 104 is further configured to replay all the stored incoming messages since the latest check point, in the order that they were stored in.

[0100] Optionally, the first controller 104 is further configured to determine that a point of failure has been reached while redoing local instructions and in response thereto cause the second controller 108 to rewind the second process 112 back to the latest checkpoint which serves as a global recovery point. Optionally, the first controller 104 and the second controller 106 are configured to perform the checkpoint in a gradual and exclusive manner, such that two processes are not performing the checkpoint at a same time.

[0101] Optionally, the first controller 104 is further configured to perform the checkpoint for the first process 110 at a first time and the second controller 108 is configured to perform the checkpoint for the second process 112 at a second time, the first time is different from the second time.

[0102] Optionally, the first time and the second time do not overlap one another.

[0103] Optionally, the method 100 further includes an orchestrator, the orchestrator is configured to indicate to a process when to perform the checkpoint.

[0104] Optionally, the orchestrator is configured to indicate to a process when to perform the checkpoint by assigning the first time and the second time.

[0105] Optionally, the orchestrator is further configured to assign the first time and the second time in order of process identifier for the first process 110 and second process 112.

[0106] Optionally, the orchestrator is configured to assign the first time and the second time based on a number of operations performed by each of the first process 110 and second process 112.

[0107] Optionally, the orchestrator is configured to assign the first time and the second time based on a number of memory operations performed by each of the first process 110 and second process 112.

[0108] Optionally, the orchestrator is configured to assign the first time and the second time based on the volume of checkpoint data stored by each of the first process 110 and second process 112.

[0109] FIG. 2 illustrates a block diagram of a first arrangement of computing resources 202 including a first controller 204 configured to execute a first process 210 and a second set of computing resources 206 including a second controller 208 configured to execute a second process 212 in accordance with an implementation of the disclosure. The first controller 204 is configured to perform a checkpoint for the first process 210. The first controller 204 is configured to store incoming messages to the first process 210. The first controller 204 is configured to determine that a failure has occurred and in response thereto, the first controller 204 is configured to reload the checkpoint for the first process 210. The first controller 204 is configured to redo local instructions starting from the checkpoint. The first controller 204 is configured to replay the stored incoming messages.

[0110] The first arrangement of computing resources 202 and the second set of computing resources 206 offer a resource-efficient, gradual, and asynchronous checkpointing method for distributed applications. The asynchronous checkpointing method of the first arrangement of computing resources 202 and the second set of computing resources 206 effectively alleviates the input / output I / O, network, and storage burdens associated with synchronous checkpointing through a gradual method that leverages message passing interface, MPI, and asynchronous I / O capabilities in order to orchestrate checkpoints among individual processes. Moreover, the asynchronous checkpointing method enables the independent self-sufficient recovery of failed processes in single-program, multiple-data, SPMD distributed applications. A combined approach of checkpointing and incoming MPI message logging ensures comprehensive self-sufficient recovery capabilities throughout the SPMD application's entire lifecycle, where the failed processes can autonomously 'recover by themselves' after a failure without necessitating the involvement of other processes. During recovery, the failed process only needs access to a latest local checkpoint and an incoming message log, as all non-failed processes are suspended and the corresponding resources may be freed for use by other jobs, thus being self-sufficient. Furthermore, the first arrangement of computing resources 202 and the second set of computing resources 206 enable the avoidance of logging sent messages, substantially reducing message logging overhead. Most importantly, the first arrangement of computing resources 202 and the second set of computing resources 206 facilitates the temporary release of resources allocated to unaffected processes until the failed process is back in operation. The first arrangement of computing resources 202 and the second set of computing resources 206 enable a resilient approach for demanding applications like high-performance computing, HPC and artificial intelligence, Al workloads. The first arrangement of computing resources 202 and the second set of computing resources 206 enable the reduction of recovery time and necessary resources through self-sufficient recovery. The first arrangement of computing resources 202 and the second set of computing resources 206 minimize message logging overhead by exclusively recording incoming messages while reproducing sent messages as outcomes of local computations. The first arrangement of computing resources 202 and the second set of computing resources 206 may be applied in the distributed training of Al and machine learning, ML models, involving the utilization of multiple worker nodes to train the model on distinct data sets, a process known as data parallel training.

[0111] The first arrangement of computing resources 202 and the second set of computing resources 206 elevate the reliability, fault tolerance, job resilience, resource efficiency, cost savings, scientific progress, user experience, and overall performance of the resource-intensive tasks. Hence, the process system ensures that these applications operate smoothly even in challenging conditions or when encountering unexpected errors.

[0112] FIG. 3 illustrates a block diagram of a second arrangement of computing resources 302 including a second controller 304 configured to execute a second process 306 in accordance with an implementation of the disclosure. The second controller 304 is configured to determine that a failure has occurred for a first process and in response thereto the second controller 304 is configured to suspend the second process 306 thereby freeing up the second set of computing resources while the first process recovers.

[0113] The second arrangement of computing resources 302 ensures comprehensive self-sufficient recovery capabilities throughout the SPMD application's entire lifecycle. This means that failed processes can autonomously 'recover by themselves' after a failure without necessitating the involvement of other processes. During recovery, each failed process only needs access to the latest local checkpoint and the incoming message log. As a result, the second arrangement of computing resources 302 not only ensures swift recovery but also significantly reduces the overhead from message logging. The second arrangement of computing resources allows the avoidance of logging sent messages, thus improving the overall system efficiency. Most importantly, the second arrangement of computing resources 302 facilitates the temporary release of resources allocated to unaffected processes until the failed process is back in operation. The key benefits of the second arrangement of computing resources 302 encompass enabling a resilient approach for demanding applications like HPC and Al workloads, the reduction of recovery time and necessary resources through self-sufficient recovery, and the minimization of message logging overhead by exclusively recording incoming messages while reproducing sent messages as outcomes of local computations.

[0114] The second arrangement of computing resources 302 elevates the reliability, fault tolerance, job resilience, resource efficiency, cost savings, scientific progress, user experience, and overall performance of these resource-intensive tasks. Hence, the second arrangement of computing resources 302 ensures that these applications operate smoothly even in challenging conditions or when encountering unexpected errors. One of the standout features of the second arrangement of computing resources 302 is the complete self-sufficient recovery capability, guaranteeing that failed processes can autonomously recover after a failure without requiring the involvement of other processes. During recovery, each failed process only needs access to the latest local checkpoint and the incoming message log. As a result, the second arrangement of computing resources 302 not only ensures swift recovery but also significantly reduces the overhead from message logging. The second arrangement of computing resources 302 allows the avoidance of logging sent messages, thus improving the overall system efficiency. To further assess the system's efficiency, various checks can be performed, including examining the I / O pattern throughout normal runtime and recovery of the application to see if and when message logs are created or accessed, verifying if only incoming messages are logged by comparing the message log size to the volume of outgoing and incoming communication traffic and forcing a process to fail to determine whether the remaining processes are suspended or remain active to assist in the bring-up.

[0115] The operations of the second arrangement of computing resources 302 enhance resilience to failure scenarios. The comprehensive approach of the second arrangement of computing resources 302 optimizes resource utilization and fault tolerance, making the second arrangement of computing resources highly suitable for resource-constrained distributed environments. In summary, the second arrangement of computing resources 302 combines adaptability, self-sufficient recovery, and efficient message logging to deliver a robust solution for HPC and Al workloads.

[0116] FIG. 4 is a block diagram of a process system 400 including a first set of computing resources 402 and a second set of computing resources 406 for recovery of data in accordance with an implementation of the disclosure. The process system 400 includes the first set of computing resources 402 includes a first controller 404 that is configured to execute a first process 410 and the second set of computing resources 406 includes a second controller 408 that is configured to execute a second process 412. The first controller 404 is configured to perform a checkpoint for the first process 410. The first controller 404 is configured to store incoming messages from the second process 412 to the first process 410. The first controller 404 is configured to determine that a failure has occurred in the first process 410 and in response thereto, the first controller 404 is configured to reload the latest checkpoint data for the first process 410, redo local computations, and replay the stored incoming messages, in order to bring-up the first process 410 to the up-to-date state i.e. latest state of the second process 412. The second controller 408 is configured to execute the second process 412 including performing local computations and / or sending messages to the first process 410. The second controller 408 is configured to determine that the failure has occurred in the first process 410 and in response thereto, the second controller 408 is configured to suspend the second process 412 thereby freeing up the second set of computing resources 406.

[0117] The process system 400 offers a resource-efficient, gradual, and asynchronous checkpointing method for distributed applications. The asynchronous checkpointing method of the process system 400 effectively alleviates the input / output I / O, network, and storage burdens associated with synchronous checkpointing through a gradual method that leverages message passing interface, MPI and asynchronous I / O capabilities in order to orchestrate checkpoints among individual processes. Moreover, the asynchronous checkpointing method enables the independent self-sufficient recovery of failed processes in single-program, multiple-data, SPMD distributed applications. A combined approach of checkpointing and incoming MPI message logging ensures comprehensive self-sufficient recovery capabilities throughout the SPMD application's entire lifecycle, where the failed processes can autonomously 'recover by themselves' after a failure without necessitating the involvement of other processes. During recovery, the failed process only needs access to a latest local checkpoint and an incoming message log, as all non-failed processes are suspended and the corresponding resources may be freed for use by other jobs, thus being self-sufficient. Furthermore, the process system 400 enables the avoidance of logging sent messages, substantially reducing message logging overhead. Most importantly, the process system 400 facilitates the temporary release of resources allocated to unaffected processes until the failed process is back in operation. The process system 400 encompasses enabling a resilient approach for demanding applications like high-performance computing, HPC and artificial intelligence, Al workloads. The process system 400 enables the reduction of recovery time and necessary resources through self-sufficient recovery. The process system 400 minimizes f message logging overhead by exclusively recording incoming messages while reproducing sent messages as outcomes of local computations. The process system 400 may be applied in the distributed training of Al and machine learning, ML models, involving the utilization of multiple worker nodes to train the model on distinct data sets, a process known as data parallel training.

[0118] The process system 400 elevates the reliability, fault tolerance, job resilience, resource efficiency, cost savings, scientific progress, user experience, and overall performance of the resource-intensive tasks. Hence, the process system 400 ensures that these applications operate smoothly even in challenging conditions or when encountering unexpected errors. One of the process system's standout features is the complete self-sufficient recovery capability, guaranteeing that failed processes can autonomously recover after a failure without requiring the involvement of other processes. During recovery, each failed process only needs access to the latest local checkpoint and the incoming message log. As a result, the process system 400 not only ensures swift recovery but also significantly reduces the overhead from message logging. The process system 400 allows the avoidance of logging sent messages, thus improving the overall system efficiency. To further assess the process system's efficiency, various checks can be performed, including examining the I / O pattern throughout normal runtime and recovery of the application to see if and when message logs are created or accessed, verifying if only incoming messages are logged by comparing the message log size to the volume of outgoing and incoming communication traffic, and forcing a process to fail to determine whether the remaining processes are suspended or remain active to assist in the bring-up.

[0119] The process system 400 is resilient to failure scenarios. Furthermore, the process system 400 ensures the checkpointing is done asynchronously without disrupting the operations. The comprehensive approach of the process system 400 optimizes resource utilization and fault tolerance, making the process system 400 highly suitable for resource-constrained distributed environments. In summary, the process system 400 combines adaptability, self-sufficient recovery, and efficient message logging to deliver a robust solution for HPC and Al workloads

[0120] Optionally, the first controller 404 is further configured to store incoming messages from the second process 412 to the first process 410 received after the checkpoint.

[0121] Optionally, the first controller 404 is further configured to send a signal to the second process 412 indicating that the first process 410 has failed. The second controller 408 is further configured to determine that the failure in the first process 410 has occurred by receiving the signal that a failure has occurred in the first process 410.

[0122] Optionally, the first controller 404 is further configured to send a second signal to the second process 412 indicating the successful recovery of the first process 410. The second controller 408 is further configured to receive the second signal indicating the successful recovery of the first process 410, and in response thereto wake the second process 412, reclaim required resources, and continue its normal execution.

[0123] Optionally, the first controller 404 is further configured to abstain from sending outgoing messages while redoing local instructions during the recovery stage.

[0124] Optionally, the first controller 404 is further configured to replay all the stored incoming messages since the latest check point, in the order that they were stored in.

[0125] Optionally, the first controller 404 is further configured to determine that a point of failure has been reached while redoing local instructions and in response thereto cause the second controller 408 to rewind the second process 412 back to the latest checkpoint which serves as a global recovery point.

[0126] FIG. 5 is a block diagram of a process system 500 for checkpointing of local data using a first set of computing resources 502 and a second set of computing resources 506 in accordance with an implementation of the disclosure. The process system 500 includes the first set of computing resources 502 including a first controller 504 configured to execute a first process 510 and the second set of computing resources 506 including a second controller 508 configured to execute a second process 512. The first controller 504 and second controller 508 are configured to perform checkpoints for the first process 510 and second process 512 in a gradual and exclusive manner, such that two processes are not performing a checkpoint at a same time.

[0127] The process system 500 offers a resource-efficient, gradual, and asynchronous checkpointing method for distributed applications. The asynchronous checkpointing method of the process system 500 effectively alleviates the input / output I / O, network, and storage burdens associated with synchronous checkpointing through a gradual method that leverages message passing interface, MPI and asynchronous I / O capabilities in order to orchestrate checkpoints among individual processes. Moreover, the asynchronous checkpointing method enables the independent self-sufficient recovery of failed processes in single-program, multiple-data, SPMD distributed applications. A combined approach of checkpointing and incoming MPI message logging ensures comprehensive self-sufficient recovery capabilities throughout the SPMD application's entire lifecycle, where the failed processes can autonomously 'recover by themselves' after a failure without necessitating the involvement of other processes. During recovery, the failed process only needs access to latest local checkpoint and an incoming message log, as all non-failed processes are suspended and the corresponding resources may be freed for use by other jobs, thus being self-sufficient. Furthermore, the process system 500 enables the avoidance of logging sent messages, substantially reducing message logging overhead. Most importantly, the process system 500 facilitates the temporary release of resources allocated to unaffected processes until the failed process is back in operation. The process system 500 encompasses enabling a resilient approach for demanding applications like high-performance computing, HPC and artificial intelligence, Al workloads. The process system 500 enables the reduction of recovery time and necessary resources through self-sufficient recovery. The process system 500 minimizes f message logging overhead by exclusively recording incoming messages while reproducing sent messages as outcomes of local computations. The process system 500 may be applied in the distributed training of Al and machine learning, ML models, involving the utilization of multiple worker nodes to train the model on distinct data sets, a process known as data parallel training.

[0128] The process system 500 elevates the reliability, fault tolerance, job resilience, resource efficiency, cost savings, scientific progress, user experience, and overall performance of the resource-intensive tasks. Hence, the process system 500 ensures that these applications operate smoothly even in challenging conditions or when encountering unexpected errors. One of the process system's standout features is the complete self-sufficient recovery capability, guaranteeing that failed processes can autonomously recover after a failure without requiring the involvement of other processes. During recovery, each failed process only needs access to the latest local checkpoint and the incoming message log. As a result, the process system 500 not only ensures swift recovery but also significantly reduces the overhead from message logging. The process system 500 allows the avoidance of logging sent messages, thus improving the overall system efficiency. To further assess the process system's efficiency, various checks can be performed, including examining the I / O pattern throughout normal runtime and recovery of the application to see if and when message logs are created or accessed, verifying if only incoming messages are logged by comparing the message log size to the volume of outgoing and incoming communication traffic, and forcing a process to fail to determine whether the remaining processes are suspended or remain active to assist in the bring-up.

[0129] The process system 500 is resilient to failure scenarios. Furthermore, the process system 500 ensures the checkpointing is done asynchronously without disrupting the operations. The comprehensive approach of the process system 500 optimizes resource utilization and fault tolerance, making the process system 500 highly suitable for resource-constrained distributed environments. In summary, the process system 500 combines adaptability, self-sufficient recovery, and efficient message logging to deliver a robust solution for HPC and Al workloads.

[0130] Optionally, the first controller 504 is further configured to perform the checkpoint for the first process 510 at a first time and the second controller 508 is configured to perform the checkpoint for the second process 512 at a second time, the first time is different from the second time.

[0131] Optionally, the first time and the second time do not overlap one another.

[0132] Optionally, the process system 500 further includes an orchestrator, the orchestrator is configured to indicate to a process when to perform the checkpoint. Optionally, the orchestrator is configured to indicate to a process when to perform the checkpoint by assigning the first time and the second time.

[0133] Optionally, the orchestrator is further configured to assign the first time and the second time in order of process identifier for the first process 510 and second process 512.

[0134] Optionally, the orchestrator is configured to assign the first time and the second time based on a number of operations performed by each of the first process 510 and the second process 512.

[0135] Optionally, the orchestrator is configured to assign the first time and the second time based on a number of memory operations performed by each of the first process 510 and the second process 512.

[0136] Optionally, the orchestrator is configured to assign the first time and the second time based on the volume of checkpoint data stored by each of the first process 510 and second process 512.

[0137] Optionally, the process system 500 is part of a High-Performance Computing, HPC, system.

[0138] Optionally, the process system 500 is configured to operate utilizing the Message Passing Interface, MPI.

[0139] FIG. 6 illustrates an exemplary representation of checkpointing and local recovery of data in accordance with an implementation of the disclosure. The exemplary representation includes a first process 602, and a second process 604. A checkpoint is performed for the first process 602 asynchronously with the second process 604. The incoming messages from the second process 604 are stored in the first process 602. A failure that has occurred in the first process 602 is determined by the second process 604 and in response thereto, the latest checkpoint data is reloaded for the first process 602. Local computations are redone, and the stored incoming messages are replayed. The second process 604 is executed to perform local computations and / or send messages to the first process 602. The failure that has occurred in the first process 602 is determined by the second process 604 and in response thereto, the second process 602 is suspended thereby freeing up the second set of computing resources.

[0140] The checkpoint for the first process 602 is performed a first time and the checkpoint for the second process 604 is performed a second time, and the first time is different from the second time.

[0141] The exemplary representation depicts asynchronous checkpoints that are performed for the first process 602 and second process 604 in a gradual and exclusive manner, such that the two processes are not performing a checkpoint at the same time. The exemplary representation depicts gradual checkpointing that utilizes Message Passing Interface, MPI, and asynchronous Input / Output capabilities in order to synchronize checkpoints of the first process 602 and the second process 604.

[0142] The MPI communication enables self-sufficient recovery of the failed processes to the latest global state of the distributed application.

[0143] A signal is sent to the second process 604 indicating that the first process 602 has failed.

[0144] A second signal to the second process 604 is sent indicating the successful recovery of the first process 602. The second signal is received by the second process 604 indicating the successful recovery of the first process 602, and in response thereto, the second process 604 is woke up, to reclaim required resources, and continues its normal execution.

[0145] When a point of failure is determined while redoing local instructions and in response thereto, the second process 604 is rewind back to the latest checkpoint which serves as a global recovery point. FIGS. 7A-7C are flow diagrams that illustrate a method for a process system to perform a checkpoint and recovery of data using a first set of computing resources and a second set of computing resources in accordance with an implementation of the disclosure. The first set of computing resources includes a first controller that is configured to execute a first process, and the second set of computing resources includes a second controller that is configured to execute a second process. At a step 702, the method includes performing a checkpoint for the first process asynchronously with the second process using the first controller. At a step 704, the method includes storing incoming messages from the second process to the first process using the first controller. At a step 706, the method includes determining that a failure has occurred in the first process using the first controller and in response thereto. At a step 708, the method includes reloading the latest checkpoint data for the first process using the first controller. At a step 710, the method includes redoing local computations, and replay the stored incoming messages using the first controller. At a step 712, the method includes executing the second process including performing local computations and / or sending messages to the first process using the second controller. At a step 714, the method includes determining that the failure has occurred in the first process using the second controller and in response thereto. At a step 716, the method includes suspending the second process thereby freeing up the second set of computing resources using the second controller. The method includes the first and second controllers performing asynchronous checkpoints for the first and second processes in a gradual and exclusive manner, such that two processes are not performing a checkpoint at a same time.

[0146] The method offers a resource-efficient, gradual, and asynchronous checkpointing method for distributed applications. The asynchronous checkpointing method of the method effectively alleviates the input / output I / O, network, and storage burdens associated with synchronous checkpointing through a gradual method that leverages message passing interface, MPI and asynchronous I / O capabilities in order to orchestrate checkpoints among individual processes. Moreover, the asynchronous checkpointing method enables the independent self-sufficient recovery of failed processes in single-program, multiple-data distributed, SPMD applications. A combined approach of checkpointing and incoming MPI message logging ensures comprehensive self-sufficient recovery capabilities throughout the SPMD application's entire lifecycle, where the failed processes can autonomously 'recover by themselves' after a failure without necessitating the involvement of other processes. During recovery, the failed process only needs access to a latest local checkpoint and an incoming message log, as all nonfailed processes are suspended and the corresponding resources may be freed for use by other jobs, thus being self-sufficient. Furthermore, the method enables the avoidance of logging sent messages, substantially reducing message logging overhead. Most importantly, the method facilitates the temporary release of resources allocated to unaffected processes until the failed process is back in operation. The method encompasses enabling a resilient approach for demanding applications like high- performance computing, HPC and artificial intelligence, Al workloads. The method enables the reduction of recovery time and necessary resources through self-sufficient recovery. The method minimizes message logging overhead by exclusively recording incoming messages while reproducing sent messages as outcomes of local computations. The method may be applied in the distributed training of Al and machine learning, ML models, involving the utilization of multiple worker nodes to train the model on distinct data sets, a process known as data parallel training.

[0147] The method elevates the reliability, fault tolerance, job resilience, resource efficiency, cost savings, scientific progress, user experience, and overall performance of the resource-intensive tasks. Hence, the method ensures that these applications operate smoothly even in challenging conditions or when encountering unexpected errors. One of the method's standout features is the complete self-sufficient recovery capability, guaranteeing that failed processes can autonomously recover after a failure without requiring the involvement of other processes. During recovery, each failed process only needs access to the latest local checkpoint and the incoming message log. As a result, the method not only ensures swift recovery but also significantly reduces the overhead from message logging. The method allows the avoidance of logging sent messages, thus improving the overall system efficiency. To further assess the method's efficiency, various checks can be performed, including examining the I / O pattern throughout normal runtime and recovery of the application to see if and when message logs are created or accessed, verifying if only incoming messages are logged by comparing the message log size to the volume of outgoing and incoming communication traffic, and forcing a process to fail to determine whether the remaining processes are suspended or remain active to assist in the bring-up.

[0148] The method is resilient to failure scenarios. Furthermore, the method ensures the checkpointing is done asynchronously without disrupting the operations. The comprehensive approach of the method optimizes resource utilization and fault tolerance, making the method highly suitable for resource-constrained distributed environments. In summary, the method combines adaptability, self-sufficient recovery, and efficient message logging to deliver a robust solution for HPC and Al workloads.

[0149] FIGS. 8A-8C are flow diagrams that illustrate a method for a method including a first set of computing resources including a first controller configured to execute a first process and a second set of computing resources including a second controller configured to execute a second process in accordance with an implementation of the disclosure. At a step 802, the method includes performing a checkpoint for the first process c. At a step 804, the method includes storing incoming messages from the second process to the first process the first controller. At a step 806, the method includes determining that a failure has occurred in the first process the first controller and in response thereto. At a step 808, the method includes reloading the latest checkpoint data for the first process the first controller. At a step 810, the method includes redoing local computations, and replaying the stored incoming messages the first controller. At a step 812, the method includes executing the second process including performing local computations and / or sending messages to the first process using the second controller. At a step 814, the method includes determining that the failure has occurred in the first process using the second controller and in response thereto. At a step 816, the method includes suspending the second process thereby freeing up the second set of computing resources using the second controller.

[0150] The method offers a resource-efficient, gradual, and asynchronous checkpointing method for distributed applications. The asynchronous checkpointing method of the method effectively alleviates the input / output I / O, network, and storage burdens associated with synchronous checkpointing through a gradual method that leverages message passing interface, MPI and asynchronous I / O capabilities in order to orchestrate checkpoints among individual processes. Moreover, the asynchronous checkpointing method enables the independent self-sufficient recovery of failed processes in single-program, multiple-data distributed, SPMD applications. Combined approach of checkpointing and incoming MPI message logging ensures comprehensive self-sufficient recovery capabilities throughout the SPMD application's entire lifecycle, where the failed processes can autonomously 'recover by themselves' after a failure without necessitating the involvement of other processes. During recovery, the failed process only needs access to latest local checkpoint and an incoming message log, as all non-failed processes are suspended and the corresponding resources may be freed for use by other jobs, thus being self-sufficient. Furthermore, the method enables the avoidance of logging sent messages, substantially reducing message logging overhead. Most importantly, the method facilitates the temporary release of resources allocated to unaffected processes until the failed process is back in operation. The method encompasses enabling a resilient approach for demanding applications like high- performance computing, HPC and artificial intelligence, Al workloads. The method enables the reduction of recovery time and necessary resources through self-sufficient recovery. The method minimizes message logging overhead by exclusively recording incoming messages while reproducing sent messages as outcomes of local computations. The method may be applied in the distributed training of Al and machine learning, ML models, involving the utilization of multiple worker nodes to train the model on distinct data sets, a process known as data parallel training.

[0151] The method elevates the reliability, fault tolerance, job resilience, resource efficiency, cost savings, scientific progress, user experience, and overall performance of the resource-intensive tasks. Hence, the method ensures that these applications operate smoothly even in challenging conditions or when encountering unexpected errors. One of the method's standout features is the complete self-sufficient recovery capability, guaranteeing that failed processes can autonomously recover after a failure without requiring the involvement of other processes. During recovery, each failed process only needs access to the latest local checkpoint and the incoming message log. As a result, the method not only ensures swift recovery but also significantly reduces the overhead from message logging. The method allows the avoidance of logging sent messages, thus improving the overall system efficiency. To further assess the method's efficiency, various checks can be performed, including examining the I / O pattern throughout normal runtime and recovery of the application to see if and when message logs are created or accessed, verifying if only incoming messages are logged by comparing the message log size to the volume of outgoing and incoming communication traffic, and forcing a process to fail to determine whether the remaining processes are suspended or remain active to assist in the bring-up.

[0152] The method is resilient to failure scenarios. Furthermore, the method ensures the checkpointing is done asynchronously without disrupting the operations. The comprehensive approach of the method optimizes resource utilization and fault tolerance, making the method highly suitable for resource-constrained distributed environments. In summary, the method combines adaptability, self-sufficient recovery, and efficient message logging to deliver a robust solution for HPC and Al workloads.

[0153] FIG. 9 is a flow diagram that illustrates a method for a method including a first set of computing resources including a first controller configured to execute a first process and a second set of computing resources including a second controller configured to execute a second process in accordance with an implementation of the disclosure. At a step 902, the method includes performing checkpoints for the first process using the first controller and the second process using the second controller in a gradual and exclusive manner, such that two processes are not performing a checkpoint at a same time.

[0154] The method offers a resource-efficient, gradual, and asynchronous checkpointing method for distributed applications. The asynchronous checkpointing method of the method effectively alleviates the input / output I / O, network, and storage burdens associated with synchronous checkpointing through a gradual method that leverages message passing interface, MPI and asynchronous I / O capabilities in order to orchestrate checkpoints among individual processes. Moreover, the asynchronous checkpointing method enables the independent self-sufficient recovery of failed processes in single-program, multiple-data distributed, SPMD applications. A combined approach of checkpointing and incoming MPI message logging ensures comprehensive self-sufficient recovery capabilities throughout the SPMD application's entire lifecycle, where the failed processes can autonomously 'recover by themselves' after a failure without necessitating the involvement of other processes. During recovery, the failed process only needs access to latest local checkpoint and an incoming message log, as all non-failed processes are suspended and the corresponding resources may be freed for use by other jobs, thus being self-sufficient. Furthermore, the method enables the avoidance of logging sent messages, substantially reducing message logging overhead. Most importantly, the method facilitates the temporary release of resources allocated to unaffected processes until the failed process is back in operation. The method encompasses enabling a resilient approach for demanding applications like high- performance computing, HPC and artificial intelligence, Al workloads. The method enables the reduction of recovery time and necessary resources through self-sufficient recovery. The method minimizes message logging overhead by exclusively recording incoming messages while reproducing sent messages as outcomes of local computations. The method may be applied in the distributed training of Al and machine learning, ML models, involving the utilization of multiple worker nodes to train the model on distinct data sets, a process known as data parallel training.

[0155] The method elevates the reliability, fault tolerance, job resilience, resource efficiency, cost savings, scientific progress, user experience, and overall performance of the resource-intensive tasks. Hence, the method ensures that these applications operate smoothly even in challenging conditions or when encountering unexpected errors. One of the method's standout features is the complete self-sufficient recovery capability, guaranteeing that failed processes can autonomously recover after a failure without requiring the involvement of other processes. During recovery, each failed process only needs access to the latest local checkpoint and the incoming message log. As a result, the method not only ensures swift recovery but also significantly reduces the overhead from message logging. The method allows the avoidance of logging sent messages, thus improving the overall system efficiency. To further assess the method's efficiency, various checks can be performed, including examining the I / O pattern throughout normal runtime and recovery of the application to see if and when message logs are created or accessed, verifying if only incoming messages are logged by comparing the message log size to the volume of outgoing and incoming communication traffic, and forcing a process to fail to determine whether the remaining processes are suspended or remain active to assist in the bring-up.

[0156] The method is resilient to failure scenarios. Furthermore, the method ensures the checkpointing is done asynchronously without disrupting the operations. The comprehensive approach of the method optimizes resource utilization and fault tolerance, making the method highly suitable for resource-constrained distributed environments. In summary, the method combines adaptability, self-sufficient recovery, and efficient message logging to deliver a robust solution for HPC and Al workloads.

[0157] In another implementation, a computer program product comprising program instructions for performing the method, when executed by one or more processors in a method.

[0158] FIG. 10 is an illustration of a computer system (e. g . , a first controller, and a second controller) in which the various architectures and functionalities of the various previous implementations may be implemented. As shown, the computer system 1000 includes at least one processor 1008 that is connected to a bus 1002, wherein the computer system 1000 may be implemented using any suitable protocol, such as Peripheral Component Interconnect, PCI-Express, Accelerated Graphics Port, AGP, Hyper Transport, or any other bus or point-to-point communication protocol. The computer system 1000 also includes a memory 1006.

[0159] Control logic (software) and data are stored in the memory 1006 which may take a form of random-access memory, RAM. In the disclosure, a single semiconductor platform may refer to a sole unitary semiconductor-based integrated circuit or chip. It should be noted that the term single semiconductor platform may also refer to multi-chip modules with increased connectivity which simulate on-chip modules with increased connectivity which simulate on-chip operation, and make substantial improvements over utilizing a conventional central processing unit, CPU and bus implementation. Of course, the various modules may also be situated separately or in various combinations of semiconductor platforms per the desires of the user.

[0160] The computer system 1000 may also include a secondary storage 1010. The secondary storage 1010 includes, for example, a hard disk drive and a removable storage drive, representing a floppy disk drive, a magnetic tape drive, a compact disk drive, digital versatile disk, DVD drive, recording device, universal serial bus, USB flash memory. The removable storage drive at least one of reads from and writes to a removable storage unit in a well-known manner.

[0161] Computer programs, or computer control logic algorithms, may be stored in at least one of the memory 1006 and the secondary storage 1010. Such computer programs, when executed, enable the computer system 1000 to perform various functions as described in the foregoing. The memory 1006, the secondary storage 1010, and any other storage are possible examples of computer-readable media.

[0162] In an implementation, the architectures and functionalities depicted in the various previous figures may be implemented in the context of the processor 1004, a graphics processor coupled to a communication interface 1012, an integrated circuit (not shown) that is capable of at least a portion of the capabilities of both the processor 1004 and a graphics processor, a chipset (namely, a group of integrated circuits designed to work and sold as a unit for performing related functions, and so forth).

[0163] Furthermore, the architectures and functionalities depicted in the various previous-described figures may be implemented in a context of a general computer system, a circuit board system, a game console system dedicated for entertainment purposes, an application-specific system. For example, the computer system 1000 may take the form of a desktop computer, a laptop computer, a server, a workstation, a game console, an embedded system. Furthermore, the computer system 1000 may take the form of various other devices including, but not limited to a personal digital assistant, PDA device, a mobile phone device, a smart phone, a television, and so forth. Additionally, although not shown, the computer system 1000 may be coupled to a network (for example, a telecommunications network, a local area network, LAN, a wireless network, a wide area network, WAN such as the Internet, a peer-to-peer network, a cable network, or the like) for communication purposes through an I / O interface 1008.

[0164] It should be understood that the arrangement of components illustrated in the figures described are exemplary and that other arrangement may be possible. It should also be understood that the various system components (and means) defined by the claims, described below, and illustrated in the various block diagrams represent components in some systems configured according to the subject matter disclosed herein. For example, one or more of these system components (and means) may be realized, in whole or in part, by at least some of the components illustrated in the arrangements illustrated in the described figures.

[0165] In addition, while at least one of these components are implemented at least partially as an electronic hardware component, and therefore constitutes a machine, the other components may be implemented in software that when included in an execution environment constitutes a machine, hardware, or a combination of software and hardware. Although the disclosure and its advantages have been described in detail, it should be understood that various changes, substitutions, and alterations can be made herein without departing from the spirit and scope of the disclosure as defined by the appended claims.

Claims

CLAIMS1. A process system (100, 400, 500), comprising a first set of computing resources (102, 402, 502) comprising a first controller (104, 204, 404, 504) configured to execute a first process (110, 210, 410, 510, 602) and a second set of computing resources (106, 206, 406, 506) comprising a second controller (108, 208, 304, 408, 508) configured to execute a second process (112, 212, 306, 412, 512, 604), wherein the first controller (104, 204, 404, 504) is configured to: perform a checkpoint for the first process (110, 210, 410, 510, 602) asynchronously with the second process (112, 212, 306, 412, 512, 604) store incoming messages from the second process (112, 212, 306, 412, 512) to the first process (110, 210, 410, 510, 602), determine that a failure has occurred in the first process (110, 210, 410, 510, 602) and in response thereto reload the latest checkpoint data for the first process (110, 210, 410, 510, 602), redo local computations, and replay the stored incoming messages, wherein the second controller (108, 208, 304, 408, 508) is configured to: execute the second process (112, 212, 306, 412, 512, 604) including performing local computations and / or sending messages to the first process (110, 210, 410, 510, 602), determine that the failure has occurred in the first process (110, 210, 410, 510, 602) and in response thereto suspend the second process (112, 212, 306, 412, 512, 604) thereby freeing up the second set of computing resources (106, 206, 406, 506).

2. The process system (100, 400, 500) according to claim 1, wherein the first controller (104, 204, 404, 504) is further configured to: store incoming messages from the second process (112, 212, 306, 412, 512, 604) to the first process (110, 210, 410, 510, 602) received after the checkpoint.

3. The process system (100, 400, 500) according to claim 1 or 2, wherein the first controller (104, 204, 404, 504) is further configured to: send a signal to the second process (112, 212, 306, 412, 512, 604) indicating that the first process (110, 210, 410, 510, 602) has failed, and wherein the second controller (108, 208, 304, 408, 508) is further configured to: determine that the failure in the first process (110, 210, 410, 510, 602) has occurred by receiving the signal that a failure has occurred in the first process (110, 210, 410, 510, 602).

4. The process system (100, 400, 500) according to claim 1, 2 or 3, wherein the first controller (104, 204, 404, 504) is further configured to: send a second signal to the second process (112, 212, 306, 412, 512, 604) indicating the successful recovery of the first process (110, 210, 410, 510), and wherein the second controller (108, 208, 304, 408, 508) is further configured to: receive the second signal indicating the successful recovery of the first process (110, 210, 410, 510, 602), and in response thereto wake the second process (112, 212, 306, 412, 512, 604), reclaim required resources, and continue its normal execution.

5. A first arrangement of computing resources (202) comprising a first controller (104, 204, 404, 504) configured to execute a first process (110, 210, 410, 510, 602) and a second set of computing resources (106, 206, 406, 506) comprising a secondcontroller (108, 208, 304, 408, 508) configured to execute a second process (112, 212, 306, 412, 512), wherein the first controller (104, 204, 404, 504) is configured to: perform a checkpoint for the first process (110, 210, 410, 510, 602), store incoming messages to the first process (110, 210, 410, 510, 602), determine that a failure has occurred and in response thereto reload the checkpoint for the first process (110, 210, 410, 510, 602), redo local instructions starting from the checkpoint, and replay the stored incoming messages.

6. A second arrangement of computing resources (302) comprising a second controller (108, 208, 304, 408, 508) configured to execute a second process (112, 212, 306, 412, 512, 604), wherein the second controller (108, 208, 304, 408, 508) is configured to: determine that a failure has occurred for a first process (110, 210, 410, 510, 602) and in response thereto suspend the second process (112, 212, 306, 412, 512, 604) thereby freeing up the second set of computing resources (106, 206, 406, 506) while the first process (110, 210, 410, 510, 602) recovers.

7. A process system (100, 400, 500), comprising a first set of computing resources (102, 402, 502) comprising a first controller (104, 204, 404, 504) configured to execute a first process (110, 210, 410, 510, 602) and a second set of computing resources (106, 206, 406, 506) comprising a second controller (108, 208, 304, 408, 508) configured to execute a second process (112, 212, 306, 412, 512, 604), wherein the first controller (104, 204, 404, 504) is configured to: perform a checkpoint for the first process (110, 210, 410, 510, 602), store incoming messages from the second process (112, 212, 306, 412, 512, 604) to the first process (110, 210, 410, 510, 602), determine that a failure has occurred in the first process (110, 210, 410, 510, 602) and in response thereto reload the latest checkpoint data for the first process (110, 210, 410, 510, 602), redo local computations, and replay the stored incoming messages, in order to bring-up the first process (110, 210, 410, 510, 602) to the up-to-date (latest) state of the second process (112, 212, 306, 412, 512, 604), wherein the second controller (108, 208, 304, 408, 508) is configured to: executing the second process (112, 212, 306, 412, 512, 604) including performing local computations and / or sending messages to the first process (110, 210, 410, 510, 602), determine that the failure has occurred in the first process (110, 210, 410, 510, 602) and in response thereto suspend the second process (112, 212, 306, 412, 512, 604) thereby freeing up the second set of computing resources (106, 206, 406, 506).

8. The process system (100, 400, 500) according to claim 7, wherein the first controller (104, 204, 404, 504) is further configured to: store incoming messages from the second process (112, 212, 306, 412, 512, 604) to the first process (110, 210, 410, 510, 602) received after the checkpoint.

9. The process system (100, 400, 500) according to claim 7 or 8, wherein the first controller (104, 204, 404, 504) is further configured to: send a signal to the second process (112, 212, 306, 412, 512, 604) indicating that the first process (110, 210, 410, 510, 602) has failed, and wherein the second controller (108, 208, 304, 408, 508) is further configured to:determine that the failure in the first process (110, 210, 410, 510, 602) has occurred by receiving the signal that a failure has occurred in the first process (110, 210, 410, 510, 602).

10. A process system (100, 400, 500), comprising a first set of computing resources (102, 402, 502) comprising a first controller (104, 204, 404, 504) configured to execute a first process (110, 210, 410, 510, 602) and a second set of computing resources (106, 206, 406, 506) comprising a second controller (108, 208, 304, 408, 508) configured to execute a second process (112, 212, 306, 412, 512, 604), wherein the first controller (104, 204, 404, 504) and second controller (108, 208, 304, 408, 508) are configured to perform checkpoints for the first controller (104, 204, 404, 504) and second controller (108, 208, 304, 408, 508) in a gradual and exclusive manner, such that two processes are not performing a checkpoint at a same time.

11. The process system (100, 400, 500) according to claim 10, wherein the first time and the second time, and do not overlap one another.

12. A method for a process system (100, 400, 500), comprising a first set of computing resources (102, 402, 502) comprising a first controller (104, 204, 404, 504) configured to execute a first process (110, 210, 410, 510, 602) and a second set of computing resources (106, 206, 406, 506) comprising a second controller (108, 208, 304, 408, 508) configured to execute a second process (112, 212, 306, 412, 512, 604), wherein the method comprises the first controller (104, 204, 404, 504) performing a checkpoint for the first process (110, 210, 410, 510, 602) asynchronously with the second process (112, 212, 306, 412, 512, 604), storing incoming messages from the second process (112, 212, 306, 412, 512, 604) to the first process (110, 210, 410, 510, 602), determining that a failure has occurred in the first process (110, 210, 410, 510, 602) and in response thereto reloading the latest checkpoint data for the first process (110, 210, 410, 510, 602), redoing local computations, and replay the stored incoming messages, wherein the method further comprises the second controller (108, 208, 304, 408, 508) executing the second process (112, 212, 306, 412, 512, 604) including performing local computations and / or sending messages to the first process (110, 210, 410, 510, 602), determining that the failure has occurred in the first process (110, 210, 410, 510, 602) and in response thereto suspending the second process (112, 212, 306, 412, 512, 604) thereby freeing up the second set of computing resources (106, 206, 406, 506), wherein the method comprises the first and second controllers performing asynchronous checkpoints for the first and second processes in a gradual and exclusive manner, such that two processes are not performing a checkpoint at a same time.

13. A method for a method (100, 400, 500), comprising a first set of computing resources (102, 402, 502) comprising a first controller (104, 204, 404, 504) configured to execute a first process (110, 210, 410, 510) and a second set of computing resources (106, 206, 406, 506) comprising a second controller (108, 208, 304, 408, 508) configured to execute a second process (112, 212, 306, 412, 512, 604), wherein the method comprises the first controller (104, 204, 404, 504) performing a checkpoint for the first process (110, 210, 410, 510, 602), storing incoming messages from the second process (112, 212, 306, 412, 512, 604) to the first process (110, 210, 410, 510, 602), determining that a failure has occurred in the first process (110, 210, 410, 510, 602) and in response thereto reloading the latest checkpoint data for the first process (110, 210, 410, 510, 602), redoing local computations, and replay the stored incoming messages, wherein the method further comprises the second controller (108, 208, 304, 408, 508)executing the second process (112, 212, 306, 412, 512, 604) including performing local computations and / or sending messages to the first process (110, 210, 410, 510, 602), determining that the failure has occurred in the first process (110, 210, 410, 510, 602) and in response thereto suspending the second process (112, 212, 306, 412, 512, 604) thereby freeing up the second set of computing resources (106, 206, 406, 506).

14. A method for a process system (100, 400, 500), comprising a first set of computing resources (102, 402, 502) comprising a first controller (104, 204, 404, 504) configured to execute a first process (110, 210, 410, 510, 602) and a second set of computing resources (106, 206, 406, 506) comprising a second controller (108, 208, 304, 408, 508) configured to execute a second process (112, 212, 306, 412, 512, 604), wherein the method comprises the first controller (104, 204, 404, 504) and second controller (108, 208, 304, 408, 508) performing checkpoints for the first process (110, 210, 410, 510, 602) and second process (112, 212, 306, 412, 512, 604) in a gradual and exclusive manner, such that two processes are not performing a checkpoint at a same time.

15. A computer program product comprising program instructions for performing the method according to claim 14, when executed by one or more processors in a process system.

Citation Information

Patent Citations

  • Backward recovery error tolerance method with forward recovery feature

    CN105242979A