Dynamically selecting a system to deploy an application based on a selected event

By acquiring user-defined information and system statistics, the system dynamically selects candidate systems for application deployment, solving the problem of insufficient consideration of multiple factors in existing disaster recovery technologies. This enables an intelligent and rapid recovery solution that supports various data protection technologies.

CN122295655APending Publication Date: 2026-06-26INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INTERNATIONAL BUSINESS MACHINE CORPORATION
Filing Date
2024-09-04
Publication Date
2026-06-26

Smart Images

  • Figure CN122295655A_ABST
    Figure CN122295655A_ABST
Patent Text Reader

Abstract

Based on the selected event, obtain user-defined information related to application recovery. Obtain data related to a set of candidate systems that can be used to deploy the application. This data includes one or more system statistics for the candidate system set, and these statistics include at least actual and historical data for one or more recovery points for the candidate system set. Perform a scoring of the candidate system set based on the user-defined information and the obtained data. Select a candidate system from the candidate system set based on this score. Initiate the deployment of the application on the selected candidate system.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] One or more aspects generally involve facilitating processing within a computing environment, and in particular, facilitating recovery within a computing environment.

[0002] Applications running on computing environments may be affected by events such as natural disasters. In these scenarios, recovery must be performed to restore application execution. During recovery, application downtime and data loss should be minimized. However, trade-offs may exist between these and other objectives, such as maintaining application performance.

[0003] Currently, various disaster recovery technologies are available, including static disaster recovery plans, load balancers, and disaster recovery orchestrators. Static disaster recovery plans are interpreted and executed by humans and take into account limited information, such as data loss. Load balancers consider some performance-related information that affects recovery time and post-recovery performance, but do not consider data loss or other considerations. Disaster recovery orchestrators consider potential data loss from configurations that include recovery point goals, but do not consider actual data loss or other system information.

[0004] Despite the existence of current disaster recovery technologies, further improvements are expected when recovering one or more applications based on the occurrence of certain events. Summary of the Invention

[0005] By providing a computer program product for facilitating processing within a computing environment, the shortcomings of the prior art are overcome, and additional advantages are provided. The computer program product includes at least one computer-readable storage medium and program instructions co-stored on the at least one computer-readable storage medium. The co-stored program instructions include program instructions for obtaining user-defined information related to the recovery of an application based on selected events, and program instructions for obtaining data related to a set of candidate systems that can be used to deploy the application. The data includes one or more system statistics for the set of candidate systems. The one or more system statistics include at least actual and historical data for one or more recovery points for the set of candidate systems. The co-stored program instructions also include program instructions for performing a scoring of the set of candidate systems based on the user-defined information and the obtained data, and program instructions for selecting a candidate system from the set of candidate systems based on the scoring. Furthermore, the co-stored program instructions include program instructions for initiating the deployment of the application on the selected candidate system.

[0006] By considering system statistics when selecting candidate systems, a comprehensive and intelligent selection / recovery solution based on multiple criteria is provided. This solution is beneficial when, for example, multiple candidate systems exist, or multiple locations including the cloud, multiple available recovery points exist. Furthermore, by considering historical data and not relying on a single point in time, the possibility of inaccuracies is reduced.

[0007] In one or more embodiments, the user-defined information includes the relative importance of at least one system statistic from one or more system statistics of one or more candidate systems in the candidate system set to the user. Considering the user-defined information enables the dynamic selection of candidate systems based on certain system statistics of the user and / or the relative importance of their objectives.

[0008] In one or more embodiments, user-defined information includes statistics of one or more user identifiers. As an example, the statistics of one or more user identifiers include an indication of a valuation function to be used to value at least one of the one or more system statistics of one or more candidate systems in the candidate system set. The use of statistics of user identifiers allows users to specify how to evaluate one or more system statistics and / or target valuations, thereby enhancing the usefulness of the post-recovery application to the user.

[0009] As an example, the one or more system statistics also include one or more resource availability metrics for the system resource set of at least one candidate system in the candidate system set. In one embodiment, the system resource set includes memory, one or more central processing units, input / output bandwidth, and latency. As another example, the one or more system statistics also include the degree of recovery for application completion. Considering resource availability metrics and / or the degree of recovery for application completion facilitates the selection of candidate systems for which the application can perform optimally.

[0010] In one or more embodiments, the program instructions for obtaining data include instructions for obtaining the data based on the occurrence of the selected event. This allows data to be obtained when the selected event occurs, and allows the latest data to be obtained and used for the selection of candidate systems.

[0011] In one or more embodiments, the candidate system set includes the current system on which the application is deployed, as well as one or more other candidate systems. Including the current system in the candidate system set improves the selection options and advantageously allows the application to remain on the same system, thereby avoiding data loss.

[0012] In one or more embodiments, the application is a stateful application, and the application data of the application is protected using one or more data protection technologies. The selection of the candidate system is independent of the one or more data protection technologies. This allows the recovery technology to be used with a wide variety of data protection technologies.

[0013] In one or more embodiments, the scoring is based on selected data related to application downtime, current resource availability of the candidate system set, time taken to restore the application, post-restore performance, and estimated post-restore resource availability. By considering multiple system statistics and / or objectives when selecting candidate systems, the speed of application restoration and the performance of the restored application are improved.

[0014] In one or more embodiments, the program instructions for performing the scoring include: program instructions for performing weighting on at least one of the one or more system statistics based on the user-defined information to obtain one or more weighted values. The program instructions for performing the scoring include program instructions for using the one or more weighted values ​​in the scoring. Weighting allows the user to input the importance of one or more system statistics, enabling the application to recover information in a user-friendly manner.

[0015] In one or more embodiments, the historical data includes one or more historical recovery time statistics for one or more candidate systems in the candidate system set. The use of historical recovery time statistics provides a more robust recovery technique, which includes a comprehensive review of information when making a selection.

[0016] In one or more embodiments, the one or more recovery point actual data includes at least one recovery point actual data for each candidate system in the candidate system set. This provides additional data for each candidate system to be considered.

[0017] In one or more embodiments, the selected event is disaster recovery from a natural disaster. A comprehensive recovery mechanism is provided for recovery from natural disasters. This mechanism is dynamic, allowing the user to input the importance of certain criteria in failover.

[0018] According to one or more aspects, each embodiment is separable and optional from each other. Furthermore, aspects of the embodiments are separable / optional from each other. Additionally, embodiments can be combined with each other.

[0019] In one aspect, a computer system is provided for facilitating processing within a computing environment. The computer system includes memory and at least one device coupled to the memory. The computer system is configured to perform a method. The method includes obtaining user-defined information related to the recovery of an application based on selected events, and obtaining data related to a set of candidate systems that can be used to deploy the application, the data including one or more system statistics for the candidate system set. The one or more system statistics include at least actual and historical data for one or more recovery points for the candidate system set. A scoring of the candidate system set is performed based on the user-defined information and the obtained data. A candidate system is selected from the candidate system set based on the scoring, and the deployment of the application on the selected candidate system is initiated.

[0020] By considering system statistics when selecting candidate systems, a comprehensive and intelligent selection / recovery solution based on multiple criteria is provided. This solution is beneficial when, for example, multiple candidate systems exist, or multiple locations including the cloud, multiple available recovery points exist. Furthermore, by considering historical data and not relying on a single point in time, the possibility of inaccuracies is reduced.

[0021] In one or more embodiments, the user-defined information includes the relative importance of at least one system statistic from one or more system statistics of one or more candidate systems in the candidate system set to the user. Considering the user-defined information enables the dynamic selection of candidate systems based on certain system statistics of the user and / or the relative importance of their objectives.

[0022] In one or more embodiments, the scoring is based on selected data related to application downtime, current resource availability of the candidate system set, time taken to restore the application, post-restore performance, and estimated post-restore resource availability. By considering multiple system statistics and / or objectives when selecting candidate systems, the speed of application restoration and the performance of the restored application are improved.

[0023] In one or more embodiments, performing a scoring process includes weighting at least one of the one or more system statistics based on the user-defined information to obtain one or more weighted values. Performing a scoring process includes using the one or more weighted values ​​in the scoring. The weighting allows the user to input the importance of one or more system statistics, enabling the application to recover information in a user-friendly manner.

[0024] In one or more embodiments, the application is a stateful application, and the application data of the application is protected using one or more data protection technologies. The selection of the candidate system is independent of the one or more data protection technologies. This allows the recovery technology to be used with a wide variety of data protection technologies.

[0025] According to one or more aspects, each embodiment is separable and optional from each other. Furthermore, aspects of the embodiments are separable / optional from each other. Additionally, embodiments can be combined with each other.

[0026] In one aspect, a computer-implemented method is provided to facilitate processing within a computing environment. This computer-implemented method includes obtaining user-defined information related to the recovery of an application based on selected events, and obtaining data related to a set of candidate systems that can be used to deploy the application. The data includes one or more system statistics for the candidate system set. The one or more system statistics include at least actual and historical data for one or more recovery points for the candidate system set. A scoring of the candidate system set is performed based on the user-defined information and the obtained data. A candidate system is selected from the candidate system set based on the scoring, and the deployment of the application on the selected candidate system is initiated.

[0027] By considering system statistics when selecting candidate systems, a comprehensive and intelligent selection / recovery solution based on multiple criteria is provided. This solution is beneficial when, for example, multiple candidate systems exist, or multiple locations including the cloud, multiple available recovery points exist. Furthermore, by considering historical data and not relying on a single point in time, the possibility of inaccuracies is reduced.

[0028] In one or more embodiments, the user-defined information includes the relative importance of at least one system statistic from one or more system statistics of one or more candidate systems in the candidate system set to the user. Considering the user-defined information enables the dynamic selection of candidate systems based on certain system statistics of the user and / or the relative importance of their objectives.

[0029] In one or more embodiments, the application is a stateful application, and the application data of the application is protected using one or more data protection technologies. The selection of the candidate system is independent of the one or more data protection technologies. This allows the recovery technology to be used with a wide variety of data protection technologies.

[0030] In one or more embodiments, the scoring is based on selected data related to application downtime, current resource availability of the candidate system set, time taken to restore the application, post-restore performance, and estimated post-restore resource availability. By considering multiple system statistics and / or objectives when selecting candidate systems, the speed of application restoration and the performance of the restored application are improved.

[0031] In one or more embodiments, performing scoring includes weighting at least one of the one or more system statistics based on the user-defined information to obtain one or more weighted values. Performing scoring includes using the one or more weighted values ​​in the scoring. Weighting allows the user to input the importance of one or more system statistics, enabling the application to recover in a user-friendly manner.

[0032] According to one or more aspects, each embodiment is separable and optional from each other. Furthermore, aspects of the embodiments are separable / optional from each other. Additionally, embodiments can be combined with each other.

[0033] In one or more aspects, a method is provided to facilitate processing within a computing environment, including computer program products, computer systems, and computer implementations. In one or more embodiments, user-defined information related to the recovery of an application is obtained based on selected events, and data related to a set of candidate systems that can be used to deploy the application is obtained. The data includes one or more system statistics for the candidate system set. The one or more system statistics include at least actual and historical data for one or more recovery points for the candidate system set. A scoring of the candidate system set is performed based on the user-defined information and the obtained data. A candidate system is selected from the candidate system set based on the scoring, and deployment of the application on the selected candidate system is initiated.

[0034] By considering system statistics when selecting candidate systems, a comprehensive and intelligent selection / recovery solution based on multiple criteria is provided. This solution is beneficial when, for example, multiple candidate systems exist, or multiple locations including the cloud, multiple available recovery points exist. Furthermore, by considering historical data, it does not rely on a single point in time, thus reducing the possibility of inaccuracies. In addition, considering user-defined information enables dynamic selection of candidate systems based on the relative importance of certain system statistics and / or objectives of the user.

[0035] In one or more aspects, a method is provided to facilitate processing within a computing environment, including computer program products, computer systems, and computer implementations. In one or more embodiments, user-defined information related to the recovery of an application is obtained based on selected events, and data related to a set of candidate systems that can be used to deploy the application is obtained. The data includes one or more system statistics for the candidate system set, and the one or more system statistics include at least actual and historical data for one or more recovery points for the candidate system set, one or more resource availability metrics for a system resource set of at least one candidate system in the candidate system set, and the degree of recovery completion for the application. A scoring of the candidate system set is performed based on the user-defined information and the obtained data. A candidate system is selected from the candidate system set based on the scoring, and deployment of the application on the selected candidate system is initiated.

[0036] By considering system statistics when selecting candidate systems, a comprehensive and intelligent selection / recovery solution based on multiple criteria is provided. This solution is beneficial when, for example, multiple candidate systems exist, or multiple locations including the cloud, multiple available recovery points exist. Furthermore, by considering historical data, it does not rely on a single point in time, thus reducing the possibility of inaccuracies. Considering resource availability metrics and / or the degree of recovery completed for the application facilitates the selection of candidate systems for which the application can perform optimally.

[0037] In one or more embodiments, the user-defined information includes the relative importance of at least one system statistic from one or more system statistics of one or more candidate systems in the candidate system set to the user. Considering the user-defined information enables the dynamic selection of candidate systems based on certain system statistics of the user and / or the relative importance of their objectives.

[0038] In one or more aspects, a method is provided for facilitating processing within a computing environment, including computer program products, computer systems, and computer implementations. In one or more embodiments, user-defined information related to the recovery of an application is obtained based on selected events, and data related to a set of candidate systems that can be used to deploy the application is obtained. The data includes one or more system statistics for the candidate system set, and the one or more system statistics include at least actual and historical data for one or more recovery points for the candidate system set. The user-defined information includes the relative importance of at least one of the one or more system statistics for one or more candidate systems in the candidate system set to the user. A scoring of the candidate system set is performed based on the user-defined information and the obtained data. A candidate system is selected from the candidate system set based on the scoring, and deployment of the application on the selected candidate system is initiated. In one example, the application is a stateful application, and the application data of the application is protected using one or more data protection technologies. The selection of the candidate system is independent of the one or more data protection technologies.

[0039] By considering system statistics when selecting candidate systems, a comprehensive and intelligent selection / recovery solution based on multiple criteria is provided. This solution is beneficial when, for example, multiple candidate systems exist, or multiple locations including the cloud, multiple available recovery points exist. Furthermore, by considering historical data, it does not rely on a single point in time, thus reducing the possibility of inaccuracies. Consideration of user-defined information enables dynamic selection of candidate systems based on the relative importance of certain system statistics and / or objectives of the user. The selection of candidate systems independent of data protection technologies allows recovery technologies to be used with a wide variety of data protection techniques.

[0040] In one or more aspects, a method is provided to facilitate processing within a computing environment, including computer program products, computer systems, and computer implementations. In one or more embodiments, user-defined information related to the recovery of an application is obtained based on selected events, and data related to a set of candidate systems that can be used to deploy the application is obtained. The data includes one or more system statistics for the candidate system set. The one or more system statistics include at least actual and historical data for one or more recovery points for the candidate system set. The user-defined information includes the relative importance of at least one of the one or more system statistics for one or more candidate systems in the candidate system set to the user. A scoring of the candidate system set is performed based on the user-defined information and the obtained data. The scoring is based on selected data related to application downtime, current resource availability of the candidate system set, time spent recovering the application, post-recovery performance, and estimated post-recovery resource availability. Candidate systems are selected from the candidate system set based on the scoring, and deployment of the application on the selected candidate system is initiated.

[0041] By considering system statistics when selecting candidate systems, a comprehensive and intelligent selection / recovery solution based on multiple criteria is provided. This solution is beneficial when, for example, multiple candidate systems exist, or multiple locations including the cloud, multiple available recovery points exist. Furthermore, by considering historical data, it does not rely on a single point in time, thus reducing the possibility of inaccuracies. Considering user-defined information enables dynamic selection of candidate systems based on the relative importance of certain user system statistics and / or objectives. By considering multiple system statistics and / or objectives when selecting candidate systems, the speed of application recovery and the performance of the recovered application are improved.

[0042] In one or more aspects, a method is provided to facilitate processing within a computing environment, including computer program products, computer systems, and computer implementations. In one or more embodiments, user-defined information related to the recovery of an application is obtained based on selected events, and data related to a set of candidate systems that can be used to deploy the application is obtained. The data includes one or more system statistics for the candidate system set. The one or more system statistics include at least actual and historical data for one or more recovery points for the candidate system set. The historical data includes historical recovery time statistics for one or more candidate systems in the candidate system set. The user-defined information includes the relative importance of at least one of the one or more system statistics for the one or more candidate systems in the candidate system set to a user. A scoring of the candidate system set is performed based on the user-defined information and the obtained data. A candidate system is selected from the candidate system set based on the scoring, and deployment of the application on the selected candidate system is initiated.

[0043] By considering system statistics when selecting candidate systems, a comprehensive and intelligent selection / recovery solution based on multiple criteria is provided. This solution is beneficial when, for example, multiple candidate systems exist, or multiple locations including the cloud, multiple available recovery points exist. Furthermore, by considering historical data, it avoids reliance on a single point in time, thus reducing the possibility of inaccuracies. Considering user-defined information enables dynamic selection of candidate systems based on the relative importance of certain system statistics and / or objectives of the user. The use of historical recovery time statistics provides a more robust recovery technique, which includes a comprehensive review of information when making a selection.

[0044] According to one or more aspects, each embodiment is separable and optional from each other. Furthermore, aspects of the embodiments are separable / optional from each other. Additionally, embodiments can be combined with each other.

[0045] This document describes and claims protection for computer-implemented methods, systems, and computer program products related to one or more aspects. Furthermore, this document also describes and claims protection for services related to one or more aspects.

[0046] Additional features and advantages are achieved through the techniques described herein. Other embodiments and aspects are described in detail herein and are considered part of the claimed aspects. Attached Figure Description

[0047] One or more aspects are specifically pointed out and clearly claimed by way of example in the claims at the end of the specification. The foregoing and objectives, features, and advantages of one or more aspects will become apparent from the following detailed description taken in conjunction with the accompanying drawings: Figure 1 An example of a computing environment for incorporating, executing, and / or using one or more aspects of this disclosure is described; Figure 2 Examples of fault domains according to one or more aspects of this disclosure are depicted; Figure 3 An example of using a recovery controller when performing failover recovery is shown, according to one or more aspects of this disclosure; Figure 4 One or more aspects of this disclosure are shown. Figure 1 An example of a submodule of the recovery module; and Figure 5 An example of a recovery process according to one or more aspects of this disclosure is described. Detailed Implementation

[0048] According to one or more aspects of this disclosure, a capability is provided to facilitate processing within a computing environment. In one aspect, processing is facilitated by providing a resilience capability that includes dynamically selecting a system (also referred to herein as a candidate system) based on selected events to deploy one or more applications. As an example, the selected event is disaster recovery from a natural disaster, but other selected events are also possible.

[0049] When the selected event occurs, such as disaster recovery due to a disaster (e.g., a natural disaster), application downtime and data loss will be minimized. However, trade-offs may exist between these and other objectives, such as maintaining application performance. Therefore, it is necessary to determine how to gather information to provide a solution in this scenario and how to select which candidate system to deploy the application on when multiple candidate systems are available. A detailed recovery plan that takes these trade-offs into account is provided.

[0050] In one or more aspects, a comprehensive recovery solution considers, for example, data loss, application downtime, performance-related information affecting recovery time and / or post-recovery performance, and any user-defined information; rapidly recommends action plans based on data obtained related to candidate systems where applications can be deployed (e.g., system statistics); and optionally, executes the recommendations.

[0051] While known disaster recovery techniques exist, they are not comprehensive. For example, static disaster recovery plans are interpreted and executed by humans, not computers, and take into account limited information, such as data loss. Furthermore, load balancers consider some performance-related information affecting recovery time and post-recovery performance, such as CPU and memory availability, but not, for example, input / output (I / O) bandwidth, historical recovery time, or data loss. Additionally, disaster recovery orchestrators consider potential data loss from configurations that include recovery point targets, without considering actual data loss or other system information such as CPU and memory availability, I / O bandwidth, and historical recovery time. Therefore, based on one or more aspects, a dynamic and comprehensive disaster recovery system is provided that automatically considers, for example, data loss, application downtime, performance-related information affecting recovery time and / or post-recovery performance, and any user-defined information (e.g., the relative importance of one or more system statistics for candidate systems and / or user-defined statistics, etc.).

[0052] In one or more aspects, a computer program product is provided for facilitating processing within a computing environment. The computer program product includes at least one computer-readable storage medium and program instructions co-stored on the at least one computer-readable storage medium. The co-stored program instructions include program instructions for obtaining user-defined information related to the recovery of an application based on selected events, and program instructions for obtaining data related to a set of candidate systems that can be used to deploy the application. The data includes one or more system statistics for the candidate system set, the one or more system statistics including at least actual and historical data for one or more recovery points of the candidate system set. The co-stored program instructions also include program instructions for performing a scoring of the candidate system set based on the user-defined information and the obtained data, and program instructions for selecting a candidate system from the candidate system set based on the scoring. Furthermore, the co-stored program instructions include program instructions for initiating the deployment of the application on the selected candidate system.

[0053] By considering system statistics when selecting candidate systems, a comprehensive and intelligent selection / recovery solution based on multiple criteria is provided. This solution is beneficial when, for example, multiple candidate systems exist, or multiple locations including the cloud, multiple available recovery points exist. Furthermore, by considering historical data and not relying on a single point in time, the possibility of inaccuracies is reduced.

[0054] Alternatively or concurrently, in one or more embodiments, the user-defined information includes the relative importance of at least one system statistic from one or more system statistics of one or more candidate systems in a candidate system set to the user. Considering the user-defined information enables the dynamic selection of candidate systems based on certain system statistics of the user and / or the relative importance of their objectives.

[0055] Alternatively or concurrently, in one or more embodiments, the user-defined information includes one or more user identification statistics. As an example, the one or more user identification statistics include an indication of a valuation function to be used to value at least one of the system statistics of one or more candidate systems in a set of candidate systems. The use of user-identified statistics allows the user to specify how to evaluate one or more system statistics and / or target values, thereby enhancing the usefulness of the post-recovery application to the user.

[0056] Alternatively or concurrently, as an example, the one or more system statistics further include one or more resource availability measures for the system resource set of at least one candidate system in the candidate system set. In one embodiment, the system resource set includes memory, one or more central processing units, input / output bandwidth, and latency. As another example, the one or more system statistics also include the degree of recovery for application completion. Considering resource availability measures and / or the degree of recovery for application completion facilitates the selection of candidate systems for which the application can perform optimally.

[0057] Alternatively or concurrently, in one or more embodiments, the program instructions for obtaining data include program instructions for obtaining data based on the occurrence of a selected event. This allows data to be obtained when the selected event occurs, and allows the latest data to be obtained and used for the selection of candidate systems.

[0058] Alternatively or concurrently, in one or more embodiments, the candidate system set includes the current system on which the application is deployed, as well as one or more other candidate systems. Including the current system in the candidate system set improves the selection options and advantageously allows the application to remain on the same system, thereby avoiding data loss.

[0059] Alternatively or concurrently, in one or more embodiments, the application is a stateful application and uses one or more data protection technologies to protect its application data. The selection of candidate systems is independent of one or more data protection technologies. This allows recovery techniques to be used with a wide variety of data protection technologies.

[0060] Alternatively or concurrently, in one or more embodiments, the scoring is based on selected data related to application downtime, current resource availability of the candidate system set, time taken to restore the application, post-restore performance, and estimated post-restore resource availability. By considering multiple system statistics and / or objectives when selecting candidate systems, the speed of application restoration and the performance of the restored application are improved.

[0061] Alternatively or concurrently, in one or more embodiments, the program instructions for performing scoring include instructions for weighting at least one of one or more system statistics based on user-defined information to obtain one or more weighted values. The program instructions for performing scoring also include instructions for using one or more weighted values ​​in the scoring. Weighting allows the user to input the importance of one or more system statistics, enabling the application to recover information in a user-friendly manner.

[0062] Alternatively or concurrently, in one or more embodiments, the historical data includes one or more historical recovery time statistics for one or more candidate systems in a set of candidate systems. The use of historical recovery time statistics provides a more robust recovery technique, which includes a comprehensive review of information when making a selection.

[0063] Alternatively or concurrently, in one or more embodiments, the actual recovery point data includes at least one actual recovery point data for each candidate system in the set of candidate systems. This provides additional data for each candidate system to be considered.

[0064] Alternatively or concurrently, in one or more embodiments, the selected event is disaster recovery from a natural disaster. A comprehensive recovery mechanism is provided for recovery from natural disasters. This mechanism is dynamic, allowing the user to input the importance of certain criteria in failover.

[0065] According to one or more aspects, each embodiment is separable and optional from each other. Furthermore, aspects of the embodiments are separable / optional from each other. Additionally, embodiments can be combined with each other.

[0066] In one aspect, a computer system is provided for facilitating processing within a computing environment. The computer system includes memory and at least one device coupled to the memory. The computer system is configured to perform a method. The method includes obtaining user-defined information related to the recovery of an application based on selected events, and obtaining data related to a set of candidate systems that can be used to deploy the application. The data includes one or more system statistics for the set of candidate systems. The one or more system statistics include at least actual and historical data for one or more recovery points for the set of candidate systems. A scoring of the set of candidate systems is performed based on the user-defined information and the obtained data. A candidate system is selected from the set of candidate systems based on the scoring, and the deployment of the application on the selected candidate system is initiated.

[0067] By considering system statistics when selecting candidate systems, a comprehensive and intelligent selection / recovery solution based on multiple criteria is provided. This solution is beneficial when, for example, multiple candidate systems exist, or multiple locations including the cloud, multiple available recovery points exist. Furthermore, by considering historical data and not relying on a single point in time, the possibility of inaccuracies is reduced.

[0068] Alternatively or concurrently, in one or more embodiments, the user-defined information includes the relative importance of at least one system statistic from one or more system statistics of one or more candidate systems in the candidate system set to the user. Considering user-defined information enables the dynamic selection of candidate systems based on certain system statistics of the user and / or the relative importance of their objectives.

[0069] Alternatively or concurrently, in one or more embodiments, the scoring is based on selected data related to application downtime, current resource availability of the candidate system set, time taken to restore the application, post-restore performance, and estimated post-restore resource availability. By considering multiple system statistics and / or objectives when selecting candidate systems, the speed of application restoration and the performance of the restored application are improved.

[0070] Alternatively or concurrently, in one or more embodiments, performing the scoring includes weighting at least one of the one or more system statistics based on the user-defined information to obtain one or more weighted values. Performing the scoring includes using the one or more weighted values ​​in the scoring. The weighting allows the user to input the importance of one or more system statistics, enabling the application to recover information in a user-friendly manner.

[0071] Alternatively or concurrently, in one or more embodiments, the application is a stateful application, and the application data of the application is protected using one or more data protection technologies. The selection of the candidate system is independent of the one or more data protection technologies. This allows the recovery technology to be used with a wide variety of data protection technologies.

[0072] According to one or more aspects, each embodiment is separable and optional from each other. Furthermore, aspects of the embodiments are separable / optional from each other. Additionally, embodiments can be combined with each other.

[0073] In one aspect, a computer-implemented method is provided to facilitate processing within a computing environment. This computer-implemented method includes obtaining user-defined information related to the recovery of an application based on selected events, and obtaining data related to a set of candidate systems that can be used to deploy the application. The data includes one or more system statistics for the candidate system set. The one or more system statistics include at least actual and historical data for one or more recovery points for the candidate system set. A scoring of the candidate system set is performed based on the user-defined information and the obtained data. A candidate system is selected from the candidate system set based on the scoring, and the deployment of the application on the selected candidate system is initiated.

[0074] By considering system statistics when selecting candidate systems, a comprehensive and intelligent selection / recovery solution based on multiple criteria is provided. This solution is beneficial when, for example, multiple candidate systems exist, or multiple locations including the cloud, multiple available recovery points exist. Furthermore, by considering historical data and not relying on a single point in time, the possibility of inaccuracies is reduced.

[0075] Alternatively or concurrently, in one or more embodiments, the user-defined information includes the relative importance of at least one system statistic from one or more system statistics of one or more candidate systems in the candidate system set to the user. Considering user-defined information enables the dynamic selection of candidate systems based on certain system statistics of the user and / or the relative importance of their objectives.

[0076] Alternatively or concurrently, in one or more embodiments, the application is a stateful application, and the application data of the application is protected using one or more data protection technologies. The selection of the candidate system is independent of the one or more data protection technologies. This allows the recovery technology to be used with a wide variety of data protection technologies.

[0077] Alternatively or concurrently, in one or more embodiments, the scoring is based on selected data related to application downtime, current resource availability of the candidate system set, time taken to restore the application, post-restore performance, and estimated post-restore resource availability. By considering multiple system statistics and / or objectives when selecting candidate systems, the speed of application restoration and the performance of the restored application are improved.

[0078] Alternatively or concurrently, in one or more embodiments, performing the scoring includes weighting at least one of the one or more system statistics based on the user-defined information to obtain one or more weighted values. Performing the scoring includes using the one or more weighted values ​​in the scoring. The weighting allows the user to input the importance of one or more system statistics, enabling the application to recover information in a user-friendly manner.

[0079] According to one or more aspects, each embodiment is separable and optional from each other. Furthermore, aspects of the embodiments are separable / optional from each other. Additionally, embodiments can be combined with each other.

[0080] In one or more aspects, a method is provided to facilitate processing within a computing environment, including computer program products, computer systems, and computer implementations. In one or more embodiments, user-defined information related to the recovery of an application is obtained based on selected events, and data related to a set of candidate systems that can be used to deploy the application is obtained. The data includes one or more system statistics for the candidate system set, and the one or more system statistics include at least actual and historical data for one or more recovery points for the candidate system set, one or more resource availability metrics for a system resource set of at least one candidate system in the candidate system set, and the degree of recovery completion for the application. A scoring of the candidate system set is performed based on the user-defined information and the obtained data. A candidate system is selected from the candidate system set based on the scoring, and deployment of the application on the selected candidate system is initiated.

[0081] By considering system statistics when selecting candidate systems, a comprehensive and intelligent selection / recovery solution based on multiple criteria is provided. This solution is beneficial when, for example, multiple candidate systems exist, or multiple locations including the cloud, multiple available recovery points exist. Furthermore, by considering historical data, it does not rely on a single point in time, thus reducing the possibility of inaccuracies. In addition, consideration of user-defined information enables dynamic selection of candidate systems based on the relative importance of certain system statistics and / or objectives of the user.

[0082] In one or more aspects, a method is provided to facilitate processing within a computing environment, including computer program products, computer systems, and computer implementations. In one or more embodiments, user-defined information related to application recovery is obtained based on selected events. Data related to a set of candidate systems that can be used to deploy the application is obtained. This data includes one or more system statistics for the candidate system set, and the one or more system statistics include at least one or more actual recovery point data, historical data of the candidate system set, one or more resource availability metrics for the system resource set of at least one candidate system in the candidate system set, and the degree of recovery completion for the application. A scoring of the candidate system set is performed based on the user-defined information and the obtained data. A candidate system is selected from the candidate system set based on the scoring, and deployment of the application on the selected candidate system is initiated.

[0083] By considering system statistics when selecting candidate systems, a comprehensive and intelligent selection / recovery solution based on multiple criteria is provided. This solution is beneficial when, for example, multiple candidate systems exist, or multiple locations including the cloud, multiple available recovery points exist. Furthermore, by considering historical data, it does not rely on a single point in time, thus reducing the possibility of inaccuracies. Considering resource availability metrics and / or the degree of recovery completed for the application facilitates the selection of candidate systems for optimal application performance.

[0084] Alternatively or concurrently, in one or more embodiments, the user-defined information includes the relative importance of at least one system statistic from one or more system statistics of one or more candidate systems in a candidate system set to the user. Considering the user-defined information enables the dynamic selection of candidate systems based on certain system statistics of the user and / or the relative importance of their objectives.

[0085] In one or more aspects, a method is provided to facilitate processing within a computing environment for computer program products, computer systems, and computer implementations. In one or more embodiments, user-defined information related to the recovery of an application is obtained based on selected events. Data related to a set of candidate systems that can be used to deploy the application is obtained. This data includes one or more system statistics of the candidate system set, and the one or more system statistics include at least one or more recovery point actual data and historical data of the candidate system set. The user-defined information includes the relative importance of at least one system statistic of one or more system statistics of one or more candidate systems in the candidate system set to the user. A scoring of the candidate system set is performed based on the user-defined information and the obtained data. A candidate system is selected from the candidate system set based on the scoring, and the deployment of the application on the selected candidate system is initiated. In one example, the application is a stateful application, and one or more data protection technologies are used to protect the application data of the application. The selection of the candidate system is independent of one or more data protection technologies.

[0086] By considering system statistics when selecting candidate systems, a comprehensive and intelligent selection / recovery solution based on multiple criteria is provided. This solution is beneficial when, for example, multiple candidate systems exist, or multiple locations including the cloud, multiple available recovery points exist. Furthermore, by considering historical data, it does not rely on a single point in time, thus reducing the possibility of inaccuracies. Consideration of user-defined information enables dynamic selection of candidate systems based on the relative importance of certain system statistics and / or objectives of the user. The selection of candidate systems independent of data protection technologies allows recovery technologies to be used with a wide variety of data protection techniques.

[0087] In one or more aspects, a method is provided to facilitate processing within a computing environment, including computer program products, computer systems, and computer implementations. In one or more embodiments, user-defined information related to application recovery is obtained based on selected events. Data related to a set of candidate systems that can be used to deploy the application is obtained. This data includes one or more system statistics for the set of candidate systems, and the one or more system statistics include at least actual and historical data for one or more recovery points of the candidate system set. The user-defined information includes the relative importance of at least one system statistic from one or more system statistics of one or more candidate systems in the candidate system set to the user. A scoring of the candidate system set is performed based on the user-defined information and the obtained data. This scoring is based on selected data related to application downtime, current resource availability of the candidate system set, time spent recovering the application, post-recovery performance, and estimated post-recovery resource availability. A candidate system is selected from the candidate system set based on the scoring, and application deployment is initiated on the selected candidate system.

[0088] By considering system statistics when selecting candidate systems, a comprehensive and intelligent selection / recovery solution based on multiple criteria is provided. This solution is beneficial when, for example, multiple candidate systems exist, or multiple locations including the cloud, multiple available recovery points exist. Furthermore, by considering historical data, it does not rely on a single point in time, thus reducing the possibility of inaccuracies. Considering user-defined information enables dynamic selection of candidate systems based on the relative importance of certain user system statistics and / or objectives. By considering multiple system statistics and / or objectives when selecting candidate systems, the speed of application recovery and the performance of the recovered application are improved.

[0089] In one or more aspects, a method is provided to facilitate processing within a computing environment, including computer program products, computer systems, and computer implementations. In one or more embodiments, user-defined information related to application recovery is obtained based on selected events. Data related to a set of candidate systems that can be used to deploy the application is obtained. This data includes one or more system statistics for the candidate system set, and the one or more system statistics include at least actual and historical data for one or more recovery points of the candidate system set. The historical data includes historical recovery time statistics for one or more candidate systems in the candidate system set. The user-defined information includes the relative importance of at least one system statistic from one or more system statistics of one or more candidate systems in the candidate system set to the user. A scoring of the candidate system set is performed based on the user-defined information and the obtained data. A candidate system is selected from the candidate system set based on the scoring, and deployment of the application on the selected candidate system is initiated.

[0090] By considering system statistics when selecting candidate systems, a comprehensive and intelligent selection / recovery solution based on multiple criteria is provided. This solution is beneficial when, for example, multiple candidate systems exist, or multiple locations including the cloud, multiple available recovery points exist. Furthermore, by considering historical data, it avoids reliance on a single point in time, thus reducing the possibility of inaccuracies. Considering user-defined information enables dynamic selection of candidate systems based on the relative importance of certain system statistics and / or objectives of the user. The use of historical recovery time statistics provides a more robust recovery technique, which includes a comprehensive review of information when making a selection.

[0091] According to one or more aspects, each embodiment is separable and optional from each other. Furthermore, aspects of the embodiments are separable / optional from each other. Additionally, embodiments can be combined with each other.

[0092] One or more aspects of this disclosure are incorporated into, executed by, and / or used by a computing environment. As an example, a computing environment can be of various architectures and types, including but not limited to: personal computing, client-server, distributed, virtual, simulation, partitioned, non-partitioned, cloud-based, quantum, grid, time-sharing, clustered, peer-to-peer, wearable, mobile, having one or more nodes, having one or more processors, and / or capable of executing processes (or processes) to, for example, select candidate systems to recover applications, perform recovery, and / or execute one or more other aspects of this disclosure. The aspects of this disclosure are not limited to a particular architecture or environment.

[0093] Various aspects of this disclosure are described by narrative text, flowcharts, block diagrams of computer systems, and / or block diagrams of machine logic included in embodiments of a computer program product (CPP). Regarding any flowchart, depending on the technology involved, operations may be performed in a different order than that shown in a given flowchart. For example, again according to the technology involved, two operations shown in consecutive flowchart blocks may be performed in reverse order, as a single integrated step, simultaneously, or in a manner that at least partially overlaps in time.

[0094] Computer Program Product Embodiment (“CPP Embodiment” or “CPP”) is a term used in this disclosure to describe any collection of one or more storage media (also referred to as “media”) collectively included in a collection of one or more storage devices, the collection of one or more storage devices collectively including machine-readable code corresponding to instructions and / or data for performing computer operations specified in a given CPP claim. A “storage device” is any tangible device capable of holding and storing instructions used by a computer processor. Without limitation, a computer-readable storage medium can be an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Some known types of storage devices that include these media include: magnetic disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory sticks, floppy disks, mechanical encoding devices (such as punch cards or pits / platforms formed in the main surface of the disk), or any suitable combination of the foregoing. Computer-readable storage media, as used in this disclosure, should not be construed as storing transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides, optical pulses through fiber optic cables, electrical signals transmitted through wires, and / or other transmission media. As those skilled in the art will understand, data is typically moved at certain incidental points in time during the normal operation of the storage device, such as during access, defragmentation, or garbage collection; however, this does not make the storage device transient, because the data is not transient when it is stored.

[0095] Reference Figure 1An example of a computing environment for executing, combining, and / or using one or more aspects of this disclosure is described. In one example, computing environment 100 includes an example of an environment for executing at least some computer code involved in performing the methods of the present invention, such as recovery code or module 150. In addition to block 150, computing environment 100 includes, for example, a computer 101, a wide area network (WAN) 102, an end-user equipment (EUD) 103, a remote server 104, a public cloud 105, and a private cloud 106. In this embodiment, computer 101 includes a processor set 110 (including processing circuitry 120 and a cache 121), communication infrastructure 111, volatile memory 112, persistent storage device 113 (including an operating system 122 and block 150, as described above), a peripheral device set 114 (including a user interface (UI) device set 123, storage device 124, and an Internet of Things (IoT) sensor set 125), and a network module 115. Remote server 104 includes a remote database 130. Public cloud 105 includes gateway 140, cloud coordination module 141, host physical machine set 142, virtual machine set 143, and container set 144.

[0096] Computer 101 can take the form of a desktop computer, laptop computer, tablet computer, smartphone, smartwatch or other wearable computer, mainframe computer, quantum computer, or any other form of computer or mobile device now known or to be developed in the future capable of running programs, accessing networks, or querying databases such as remote database 130. As is well known in the field of computer technology, and depending on the technology, the performance of a computer-implemented method can be distributed across multiple computers and / or multiple locations. On the other hand, in this presentation of computing environment 100, the detailed discussion focuses on a single computer, specifically computer 101, to keep the presentation as simple as possible. Computer 101 can reside in the cloud, even... Figure 1 It is not shown in the cloud, and on the other hand, computer 101 does not need to be in the cloud unless it can be indicated with certainty to any extent.

[0097] Processor assembly 110 includes one or more computer processors of any type now known or to be developed in the future. Processing circuitry 120 may be distributed across multiple packages, such as multiple cooperating integrated circuit chips. Processing circuitry 120 may implement multiple processor threads and / or multiple processor cores. Cache 121 is memory located within the processor chip package and is typically used for data or code that should be readily accessible by the threads or cores running on processor assembly 110. Cache memory is typically organized into multiple levels based on its relative proximity to the processing circuitry. Alternatively, some or all of the cache in the processor assembly may be located “off-chip.” In some computing environments, processor assembly 110 may be designed to work with qubits and perform quantum computing.

[0098] Computer-readable program instructions are typically loaded onto computer 101 to cause the processor set 110 of computer 101 to perform a series of operational steps to implement a computer-implemented method, such that the instructions thus executed instantiate the method specified in the flowcharts and / or descriptive descriptions of the computer-implemented method included in this document (collectively, the “method of the invention”). These computer-readable program instructions are stored in various types of computer-readable storage media, such as cache 121 and other storage media discussed below. The program instructions and associated data are accessed by the processor set 110 to control and direct the execution of the method of the invention. In computing environment 100, at least some of the instructions for performing the method of the invention may be stored in block 150 of permanent storage device 113, in block 400.

[0099] Communication structure 111 is a signal transmission path that allows the various components of computer 101 to communicate with each other. Typically, this structure consists of switches and conductive paths, such as switches and conductive paths that form buses, bridges, physical input / output ports, etc. Other types of signal communication paths can be used, such as fiber optic communication paths and / or wireless communication paths.

[0100] Volatile memory 112 is any type of volatile memory now known or to be developed in the future. Examples include dynamic random access memory (RAM) or static RAM. Typically, volatile memory 112 is characterized by random access, but this is not necessary unless explicitly stated otherwise. In computer 101, volatile memory 112 is located in a single package and is internal to computer 101; however, alternatively or additionally, volatile memory may be distributed across multiple packages and / or located externally relative to computer 101.

[0101] The persistent storage device 113 is any form of non-volatile memory for a computer, now known or to be developed in the future. The non-volatility of this memory means that the stored data is retained regardless of whether power is supplied to the computer 101 and / or directly to the persistent storage device 113. The persistent storage device 113 may be a read-only memory (ROM), but typically at least a portion of the persistent memory allows data to be written, deleted, and rewritten. Some common forms of persistent storage include hard disks and solid-state storage devices. The operating system 122 may take several forms, such as various known proprietary operating systems or operating systems employing an open-source portable operating system interface type with a kernel. The code included in block 150 generally includes at least some of the computer code involved in performing the methods of the present invention.

[0102] Peripheral device set 114 includes a set of peripheral devices for computer 101. Data communication connections between peripheral devices and other components of computer 101 can be implemented in various ways, such as Bluetooth connectivity, near field communication (NFC) connectivity, connections made by cables (such as Universal Serial Bus (USB) type cables), plug-in connections (e.g., secure digital (SD) cards), connections made through local area communication networks, and even connections made through wide area networks such as the Internet. In various embodiments, UI device set 123 may include components such as displays, speakers, microphones, wearable devices (such as goggles and smartwatches), keyboards, mice, printers, touchpads, game controllers, and haptic devices. Storage device 124 is an external storage device, such as an external hard drive, or a pluggable storage device, such as an SD card. Storage device 124 can be permanent and / or volatile. In some embodiments, storage device 124 may take the form of a quantum computing storage device for storing data in the form of qubits. In embodiments where computer 101 requires substantial storage (e.g., where computer 101 locally stores and manages a large database), this storage can be provided by peripheral storage devices designed for storing very large amounts of data, such as a Storage Area Network (SAN) shared by multiple geographically distributed computers. The IoT sensor set 125 comprises sensors that can be used in IoT applications. For example, one sensor could be a thermometer, while another could be a motion detector.

[0103] Network module 115 is a collection of computer software, hardware, and firmware that allows computer 101 to communicate with other computers via WAN 102. Network module 115 may include hardware such as a modem or Wi-Fi transceiver, software for packetizing and / or depacketizing data transmitted over the communication network, and / or web browser software for transmitting data over the Internet. In some embodiments, the network control and network forwarding functions of network module 115 are performed on the same physical hardware device. In other embodiments (e.g., embodiments utilizing Software-Defined Networking (SDN)), the control and forwarding functions of network module 115 are performed on physically separate devices, such that the control function manages several different network hardware devices. Computer-readable program instructions for performing the methods of the present invention can typically be downloaded to computer 101 from an external computer or external storage device via a network adapter card or network interface included in network module 115.

[0104] WAN 102 is any wide area network (e.g., the Internet) capable of transmitting computer data over non-local distances using any technology known now or developed in the future for transmitting computer data. In some embodiments, WAN 102 may be replaced by and / or supplemented by a local area network (LAN) designed to transmit data between devices located in a local area, such as a Wi-Fi network. WANs and / or LANs typically include computer hardware such as copper transmission cables, fiber optic transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and edge servers.

[0105] End User Equipment (EUD) 103 is any computer system used and controlled by an end user (e.g., a customer of the enterprise operating computer 101) and can take any of the forms discussed above in conjunction with computer 101. EUD 103 typically receives useful and available data from the operation of computer 101. For example, assuming computer 101 is designed to provide recommendations to an end user, these recommendations are typically transmitted from network module 115 of computer 101 to EUD 103 via WAN 102. In this way, EUD 103 can display or otherwise present recommendations to the end user. In some embodiments, EUD 103 can be client equipment such as a thin client, heavy client, mainframe, desktop computer, etc.

[0106] Remote server 104 is any computer system that provides at least some data and / or functionality to computer 101. Remote server 104 can be controlled and used by the same entity operating computer 101. Remote server 104 represents a machine that collects and stores useful and useful data used by other computers such as computer 101. For example, if computer 101 is designed and programmed to provide recommendations based on historical data, that historical data can be provided to computer 101 from a remote database 130 of remote server 104.

[0107] Public cloud 105 is any computer system that can be used by multiple entities, providing on-demand availability of computer system resources and / or other computing capabilities (particularly data storage (cloud storage) and computing power) without the need for direct, active management by users. Cloud computing typically leverages resource sharing to achieve scalability consistency and economy. Direct and active management of the computing resources of public cloud 105 is performed by the computer hardware and / or software of cloud coordination module 141. The computing resources provided by public cloud 105 are typically implemented by virtual computing environments running on various computers constituting the host physical machine set 142, which is the entirety of physical computers in and / or available to the public cloud 105. Virtual computing environments (VCEs) typically take the form of virtual machines from virtual machine set 143 and / or containers from container set 144. It should be understood that these VCEs can be stored as images and can be transferred between various physical machine hosts as images or after the VCEs are instantiated. Cloud coordination module 141 manages the transfer and storage of images, deploys new instantiations of VCEs, and manages the active instantiation of VCE deployments. Gateway 140 is a collection of computer software, hardware, and firmware that allow public cloud 105 to communicate via WAN 102.

[0108] Now, we will provide some further explanation of Virtualized Computing Environments (VCEs). A VCE can be stored as an "image." A new active instance of a VCE can be instantiated from this image. Two common types of VCEs are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to an operating system feature where the kernel allows multiple isolated user-space instances, called containers, to exist. From the perspective of the programs running within them, these isolated user-space instances typically appear as actual computers. Computer programs running on a regular operating system can utilize all the resources of that computer, such as connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, programs running within a container can only use the contents of the container and the devices allocated to the container; this is a characteristic known as containerization.

[0109] Private cloud 106 is similar to public cloud 105, except that computing resources are available only to a single enterprise. While private cloud 106 is depicted as communicating with WAN 102, in other embodiments, private cloud may be completely disconnected from the Internet and accessible only via a local / private network. A hybrid cloud is a combination of multiple clouds of different types (e.g., private, community, or public cloud types) typically implemented by different vendors. Each of the multiple clouds remains a separate and discrete entity, but the larger hybrid cloud architecture is bound together by norms or proprietary technologies that enable coordination, management, and / or data / application portability across the multiple component clouds. In this embodiment, public cloud 105 and private cloud 106 are both part of a larger hybrid cloud.

[0110] The computing environment described above is merely one example of a computing environment used in conjunction with, to execute, and / or use one or more aspects of this disclosure. Other examples are possible. For instance, in one or more embodiments, Figure 1 One or more components / modules / blocks are not included in the computing environment and / or are not used in one or more aspects of this disclosure. Furthermore, in one or more embodiments, additional and / or other components / modules / blocks may be used. Other variations are possible.

[0111] Depending on one or more aspects, the processors and / or machines of the computing environment may be located at different sites. Sites may be geographically proximate to one or more other sites and / or geographically distant from one or more other sites. In one or more examples, at least one site may be a cloud environment. Many examples are possible. A site may be considered a failover site if a selected event, such as disaster recovery, occurs due to a disaster at another site, and that site is used to apply failover. Disasters are, for example, natural disasters, including but not limited to earthquakes; tsunamis; floods; volcanic eruptions; landslides; fires; weather-related disasters such as hurricanes, blizzards, tornadoes, cyclones, monsoons, etc.; and / or other natural disasters. In other examples, the disaster may be other types of disasters and / or the selected event may be other types of selected events.

[0112] Figure 2 The document describes an example of multiple failover sites. In one example, multiple sites exist, including, for example, site 1 (200a), site 2 (200b), and site 3 (200c); in other examples, additional, fewer, and / or other sites may exist. Each site is considered a fault domain 210a, 210b, 210c of each other, and each fault domain is a candidate system in a set of candidate systems to be considered for deploying applications from the fault domain.

[0113] Candidate systems are, for example, any system capable of hosting applications that provide disaster recovery protection. The set of candidate systems includes, for example, the system to which the application is currently deployed, and one or more other systems that have access to the application's data (e.g., one or more volumes of the application and / or copies of those volumes). If data is replicated asynchronously, candidate systems with data copies may not have the latest data.

[0114] Each candidate system has a set of system statistics associated with it, including, for example, one or more recovery point-related objectives and one or more recovery time-related objectives. As an example, recovery point-related objectives include the Recovery Point Objective (RPO), which is the maximum tolerable time between a replica of its data and a primary replica of its data, and / or Recovery Point Actual Data (RPA), which is the time between a replica of its data and a primary replica of its data. In the case of a disaster, it is assumed that the primary replica of the data begins from the time of disaster detection. In one example, the recovery point objective indicates how much data loss can be tolerated within a selected time period. The Recovery Point Actual Data indicates the time elapsed since the last successful recovery point. It helps derive the amount of data that may be lost during failover. Recovery point-related objectives may include more, less, and / or other objectives.

[0115] As an example, recovery time-related objectives include, for instance, past or historical recovery times, current resource availability, and / or the extent to which the application has been recovered. Recovery time-related objectives may include additional, fewer, and / or other objectives.

[0116] Recovery time is, for example, the time from the start of an event (such as a disaster) to the successful completion of recovery on the system selected for recovery. In one example, the application's past recovery time is disaster-protected, making this information available in the event of a disaster.

[0117] Current resource availability helps determine whether there are sufficient resources (e.g., central processing unit (CPU), memory, disk capacity, storage, input / output (I / O) bandwidth, and / or I / O latency, etc.) to allow an application to perform well. Resources are referred to in this document as system statistics (e.g., performance-related system statistics that may affect recovery time and / or post-recovery performance). Excessive resource availability may not improve application performance. By default, system statistics have a practical upper limit. If a system's resources or performance exceed this limit, it is capped at the upper limit. For example, suppose the upper limit for CPU cores is 100. If system X has a CPU count of 128, it is scored (as described below) as if it had 100 cores. Other examples are possible. The current resource availability attribute may also consider historical values, as a single point in time can lead to inaccuracies. This is still represented as a single value for scoring purposes.

[0118] The degree of application recovery indicates the application's state during recovery. For example, when performing disaster recovery on an application on a candidate system, the application to be recovered can be in one of the following states: Cold: The application has not yet been installed or deployed; Gradually warming: Only some parts of the application are deployed but inactive; Warm: All parts of the application are installed or deployed, but inactive; Hot: All parts of the application are installed, deployed, and active. Additional, lesser, and / or other states can be indicated.

[0119] In one example, each state can be converted into a value for scoring. For example, cold: 25, warm: 50; gradually warming: 75, hot: 100. If system A is warm, the return value is 50. Additional, fewer, and / or other examples of state and / or score values ​​are possible.

[0120] The system statistics set may include additional, fewer, and / or other system statistics. For example, in one example, it includes estimated post-recovery resource availability (e.g., estimated resource availability such as estimated storage capacity, other estimated system resources, etc.).

[0121] continue Figure 2 In one example, if application (App) 220 running on site 2 (200b) fails due to an event such as a disaster or other event, App A can failover to fault domain 210A or fault domain 210c, or it can remain in fault domain 210b. The selection of fault domain (or candidate system) is based on the chosen objectives, including but not limited to minimizing application downtime, minimizing data loss, and / or maintaining application performance. Figure 2 The example described illustrates two objectives: data loss and recovery time estimation.

[0122] In one example, if application A 220 is moved to fault domain (or candidate system) 210A, data loss 250A is considered moderate, and recovery time estimate 252A is considered low; if it is moved to fault domain (or candidate system) 210c, data loss 250c is considered low, and recovery time estimate 252c is considered moderate. Furthermore, if application A 220 remains in fault domain (or candidate system) 210b, data loss 250b is zero, but recovery time estimate 252b is considered high. Therefore, in this example, data loss and recovery time estimate are considered when determining where to recover application A 220. However, depending on one or more aspects, additional, fewer, and / or other objectives may be considered, as well as one or more system statistics and / or user-defined information, as described herein.

[0123] While this article provides example values ​​for data loss and recovery time estimation, these are merely examples. Values ​​may be determined based on, for example, the chosen indication; on a range; relative to other values, etc. Additional, fewer, and / or other values ​​and / or value types may be used. Furthermore, additional, fewer, and / or other objectives and / or information may be considered. Many examples and / or variations are possible.

[0124] According to one or more aspects, a resilience capability is provided that includes dynamically selecting candidate systems from a set of candidate systems based on selected events such as disaster recovery to deploy one or more applications. Disasters can be, for example, natural disasters (e.g., earthquakes, tsunamis, floods, volcanic eruptions, landslides, fires, weather-related disasters such as hurricanes, blizzards, tornadoes, cyclones, monsoons, etc.; and / or other natural disasters) and / or other types of disasters or selected events.

[0125] Based on the occurrence of the selected event, recovery will be performed for one or more applications. During recovery, selected objectives, such as minimizing application downtime and data loss, are considered when selecting the systems to be used in the recovered applications. However, trade-offs may exist between these objectives and other objectives, such as maintaining application performance. In one or more aspects, a comprehensive solution is provided that considers (e.g., dynamically) multiple criteria, including, for example, one or more objectives; one or more system statistics; and / or user-defined information, such as user-specified objectives and / or user-identified statistics (e.g., a set of system statistics important to the user and / or other statistics (e.g., time of day, etc.)) and an indication of the relative priority or importance of these statistics to the user.

[0126] refer to Figure 3 An example is described of considering and / or using multiple criteria when selecting candidate systems to be used in a recovery application. In one example, user 310 (or one or more users) performs pre-configuration 312, where the user inputs user-defined information, such as the relative importance and / or other information of various system statistics 314 that contribute to the selection of candidate systems. System statistics include, for example, objectives related to the recovery point, objectives related to recovery time, estimated post-recovery resource availability, and / or other system statistics / statistics affecting recovery and / or post-recovery. In one example, for each statistic identified by the user (or selected by the user-identified statistic), the user can specify (e.g., via user-defined information) the evaluation function (e.g., linear, logarithmic, exponential, etc.) and / or normalization function to be used in scoring the statistic. Furthermore, in one or more examples, the user can specify one or more weights indicating the relative importance of the statistic. Other examples are possible.

[0127] Based on a selected event 350, such as recovery from a natural disaster, a recovery controller 360 is activated, which obtains (e.g., collects) data such as one or more system statistics. Using, for example, a data collector 362, one or more system statistics are collected from one or more systems 380 (e.g., System 1 (S1), System 2 (S2), System 3 (S3)) and / or one or more data repositories. The one or more system statistics include, for example, current data (e.g., current system statistics) and historical data (e.g., historical system statistics). As an example, the one or more system statistics include recovery point-related objectives (e.g., recovery point objectives, actual recovery point data, etc.); recovery time-related objectives (e.g., current resource availability; degree of application recovery completed; historical recovery time-related objectives, such as past recovery times, etc.); and / or, as an example, estimated post-recovery resource availability. The obtained system statistics may include additional, lesser, and / or other system statistics and / or other statistics.

[0128] In one example, the acquired data is forwarded to system scorer 364, which scores each system, including the current system, for failover. In another example, each candidate system is scored based on user input (e.g., user-defined information) and collected data (e.g., system statistics). Based on this score, system selector 366 selects a candidate system from the candidate system set for failover. As an example, the best candidate system is selected based on the score. The best candidate system is defined, for example, as the candidate system with the lowest (or highest) score relative to other candidate systems in the candidate system set. In other examples, the best candidate system is defined based on other indicators. For example, it can be defined as a candidate system with a predefined relationship to one or more indicators and another predefined relationship to one or more other indicators. Many examples and variations are possible.

[0129] In one example, the recovery controller 360 also includes an application (App) deployer 368 to deploy the application on a selected candidate system indicated by the system selector 366.

[0130] In one or more examples, the recovery controller executes on a system (referred to as the leader system). The leader system is, for example, one or more devices, such as one or more computers (e.g., computer 101 and / or other computers); one or more servers (e.g., remote server 104 and / or other servers); one or more end-user devices (e.g., end-user device 103 and / or other end-user devices); one or more processors or nodes (e.g., processors or nodes of processor set 110 and / or other processors or nodes); processing circuitry (e.g., processing circuitry 120 of processor set 110 and / or other processing circuitry); and / or other devices, etc. Additional and / or other computers, servers, end-user devices, processors, nodes, processing circuitry, and / or other devices may be selected as the leader system. Many examples are possible.

[0131] The leader system can be selected before or after a selection event (e.g., a natural disaster) and can be selected manually (e.g., by users, administrators, etc.) or automatically (e.g., by a process, using artificial intelligence, or machine learning, etc.). In the case of using machine learning, the model is trained to select the leader system and retrained based on current and / or learned information.

[0132] In one or more aspects, in order to perform recovery, a recovery controller such as recovery controller 360 may use a recovery module such as recovery module 150 to select candidate systems for failover. In one example, the recovery module (e.g., recovery module 150) includes various sub-modules for facilitating and / or performing recovery and / or related tasks. Sub-modules are, for example, computer-readable program code (e.g., instructions) in a computer-readable medium, such as a storage device (as an example, persistent storage device 113, cache 121, storage device 124, other storage devices). Although, as an example, recovery module 150... Figure 1 It is depicted as being in persistent storage device 113, but one or more submodules may be in other storage devices, etc. Many variations are possible.

[0133] Computer-readable media may be part of one or more computer program products, and computer-readable program code may be executed and / or used to execute by one or more devices (e.g., one or more computers, such as computer 101 and / or other computers; one or more servers, such as remote server 104 and / or other servers; one or more end-user devices, such as end-user device 103 and / or other end-user devices; one or more processors or nodes, such as processors or nodes of processor set 110 and / or other processors or nodes; processing circuitry, such as processing circuitry 120 of processor set 110 and / or other processing circuitry; and / or other devices, etc.). Additional and / or other computers, servers, end-user devices, processors, nodes, processing circuitry, and / or other devices may be used to execute one or more submodules and / or portions thereof. Numerous examples are possible.

[0134] Reference Figure 4 An example of a submodule describing recovery module 150 is provided. As an example, recovery module 150 includes, for example, a user-defined information acquisition submodule 410 for acquiring user preferences related to one or more objectives and / or system statistics; a data collection submodule 420 for collecting data (e.g., system statistics) related to candidate systems that can be used for failover; a scoring submodule 430 for scoring available candidate systems based on, for example, user-defined information and the collected data; a selection submodule 440 for selecting candidate systems to be used for failover; and a deployment submodule 450 for initiating deployment and / or deploying applications on the selected candidate systems in accordance with one or more aspects of this disclosure. Additionally, fewer and / or other submodules may be used to perform recovery and / or related tasks. Other variations are possible. Although various submodules are described, recovery modules such as recovery module 150 may include additional, fewer, and / or different submodules. Specific submodules may include additional code that includes code from other submodules, fewer codes, and / or different codes. Furthermore, additional and / or other modules may be used to facilitate recovery and / or perform related tasks. Many changes are possible.

[0135] In one example, the data collector 362, system scorer 364, system selector 366, and application deployer 368 can be implemented using one or more of submodules 420-450. In other examples, other implementation techniques can be used. Many examples are possible.

[0136] As described in this document, one or more submodules are used to select candidate systems for failover, and the application's failover is initiated on the selected candidate systems, as referred to in this document. Figure 5In one example, the recovery process (e.g., recovery process 500) is implemented using one or more submodules (e.g., one or more submodules 410-450) and executed by one or more devices (e.g., one or more computers (e.g., computer 101, other computers, etc.), one or more servers (e.g., server 104, other servers, etc.), one or more end-user devices (e.g., end-user device 103, other end-user devices), one or more processors, nodes, and / or processing circuitry (e.g., processor set 110 or other processor sets) and / or one or more other devices, etc.). Although example computers, servers, end-user devices, processors, nodes, processing circuitry, and / or devices are provided, additional, fewer, and / or other computers, servers, end-user devices, processors, nodes, processing circuitry, and / or devices may be used for recovery and / or other processing. Various options are possible.

[0137] In one example, the recovery process is performed on the leader system. Other examples are also possible.

[0138] In one example, reference Figure 5 Recovery process 500 (also referred to as process 500) includes obtaining user-defined information 510 based on pre-configuration performed by one or more users. For example, one or more users provide user-defined information such as valuation inputs (e.g., valuation functions and / or weights) for one or more selected criteria, indicating, for example, the importance of one or more criteria to the user. These criteria include, for example, system statistics, including but not limited to one or more targets related to the recovery point, one or more targets related to recovery time, estimated post-recovery resource availability and / or other system statistics; selected targets related to, for example, application downtime, application recovery time and / or post-recovery performance; and / or additional, fewer and / or other criteria, etc. As an example, these criteria are provided prior to a selected event, such as recovery from a disaster. In another example, these criteria may be provided after the event occurs but before selecting candidate systems on which applications will be recovered. Other examples are possible.

[0139] Process 500 receives an indication that an application failure and / or an event has occurred (520). As an example, process 500 receives this indication based on a failure of the application heartbeat process and / or a failure of a status query. If the heartbeat process or status query does not receive a response from its running application or system, or does not receive an indication to perform a failover, process 500 continues with recovery.

[0140] In one example, as part of the recovery process, process 500 collects 530 data, such as current and historical data related to one or more candidate systems in the candidate system set and / or the failover application. As an example, the collected data includes one or more system statistics (e.g., one or more current and / or historical system statistics) for one or more candidate systems. The collected data includes values ​​representing, for example, one or more recovery point targets, one or more actual recovery point data, one or more past recovery times, current resource availability, the extent of completed application recovery, and / or estimated post-recovery resource availability. Additionally, supplementary, less detailed, and / or other system statistics may be collected. As an example, this data (e.g., system statistics) is obtained from candidate systems and / or retrieved from memory, storage devices, and / or other accessible locations.

[0141] Based on the obtained user-defined information and collected data, process 500 scores each candidate system (or selected candidate system) 540. As an example, for each candidate system, estimation, normalization, and / or weighting are performed on the statistics of each system (or selected system statistics) 545, and the results are summed to provide a score for the candidate system. Estimation is performed using functions such as linear, logarithmic, exponential, etc., which may be provided as user preferences, as default values, or determined by processes such as machine learning. Normalization places statistical values ​​at the same level (e.g., the same type or format). Weighting factors are used in the preferences of one or more users. As an example, each preference is a value between 0 and 1; other values ​​are also possible. Further details are provided below.

[0142] An example of normalization is range normalization, where all values ​​are linearly mapped to, for example, [0, 100] after a maximum value is defined. For instance, a general formula includes, for example, for an output range [X1, X2] and an input range [X1, X2], with an input x and an output x, X = X1 + ({X2 - X1} / {X2 - X1})(x - X1). If the output range is constrained to [0, 100] and X1 = 0 is used, then a maximum range for each value is defined to map it to the chosen dimension, and X = (100 / X2) * x is used.

[0143] Example 1: The maximum core limit is 128; therefore, a system with 16 cores will have a score of X = (100 / 128) * 16 = 12.5.

[0144] Example 2: The maximum storage limit is 256GB; therefore, a system with 60 GB of free storage will have a score of X = (100 / 256) * 60 = 23.4375.

[0145] Example 3: The maximum I / O bandwidth (BW) limit is 4000 MB / s; therefore, a system with a bandwidth of 200 MB / s will have a score X = (100 / 4000) * 200 = 5. Many examples are possible. Furthermore, other normalization equations and / or techniques can be used.

[0146] In one example, one or more system statistics (and / or objectives, etc.) are multiplied by one or more weights.

[0147] In one example, weights determine the relative value of a system statistic (and / or objective, etc.). As the quantity of a particular system statistic (and / or objective, etc.) increases or decreases, its value may not change proportionally. For example: each increase in CPU can benefit the application until a cap is reached, where additional CPU provides little or no benefit; for example, a recovery point in the most recent 5 minutes may be more valuable than those in previous minutes; each system statistic (and / or objective) can be valued by a function before being multiplied by its weight. Example functions include: linear (e.g., the default): F(a) = a; linear with an upper limit (increasing to a certain limit): F(a, a) = max(a, a) ceilling Logarithm (decreasing returns): f(a) = log(a); Combination: f(a) = max(log(a), a) ceiling Additional, fewer, and / or other functions can be specified and used.

[0148] Exemplary terms used in this document may include, for example: System Statistics = a, such as: actual recovery point data, recovery point target, idle CPU, idle GB of memory, etc. Estimated System Statistics = valuate(a) - the estimate by the function. Factor = F(a) - the estimated and normalized system statistics; F(a) = normalized(valueate(a)). System Statistics Weights = w(a) - the relative values ​​of the estimated system statistics; valid values ​​range from, for example, 0 to 100%; all weights sum to the same value, for example, 100%. Weighting Factor: F weightd (a) = F(a). Other examples and variations are possible.

[0149] In one example, the scoring process transforms system statistics and / or objectives or other criteria of different units and scales into unitless values ​​and combines them to produce a single system score. An example of a scoring formula includes:

[0150] Score i =

[0151] in:

[0152] i = current system (index / iterator), n = total number of systems, j = current factor index;

[0153] Statistics: Data on actual recovery points reported by the system, resource availability, etc.; and

[0154] Weights: such as the relative importance of individual factors determined a priori by, for example, the user. In other examples, machine learning and / or artificial intelligence can be used to determine weights.

[0155] In one example, for normalization, the sum of the weights equals 1, or .

[0156] In one example, a lower score implies a more suitable candidate system; however, a score inverse process can be used for higher scores to suggest a better fit. Other examples are possible.

[0157] In one example, for some system statistics that include time measurements (e.g., actual recovery point data, I / O response time, etc.), a lower score is considered better (lower is a better unit), while for other system statistics that include resource availability (e.g., CPU cores, I / O bandwidth, etc.), a higher score is considered better (higher is a better unit). Therefore, in one example, when using a lower score to select the most suitable (e.g., best) candidate system, a "higher is-good" unit is converted to a "lower is-good" unit by subtracting the value of "higher is-good" from the maximum value. For example, assuming the maximum number of CPU cores is 100, a candidate system with 60 available cores is converted to 100-60=40, and another candidate system with 20 available cores is converted to 100-20=80. Other examples are possible.

[0158] In one example, the value of a particular statistic may not change proportionally as the quantity increases or decreases. Estimation functions are used to transform the value appropriately (e.g., with linear, exponential, logarithmic, etc.). Other examples are possible.

[0159] A specific example of rating includes:

[0160] System statistical weights input by the user: RPA weight = 0.6; waiting time weight = 0.4.

[0161] System 1 statistics: RPA = 10; Precision = 5.

[0162] Score system1 =(10*0.6)+(5*0.4)=6+2=8.

[0163] Although the scoring formulas and / or techniques described above are used, other scoring formulas and / or techniques may also be used. Furthermore, additional, fewer, and / or other systematic statistics and / or criteria may be used for weighting. Other examples are possible.

[0164] In one example, given the requirements of an application to be failoverd, process 500 filters 550 candidate systems with insufficient resources (if any). Based on the filtering, process 500 determines 560 whether any candidate systems remain with sufficient resources (e.g., CPU, memory, I / O bandwidth, and / or other selected resources) to execute the application. If no systems remain, process 500 returns to collect data on the candidate systems (e.g., system statistics). In one example, exponential backoff is used to perform retries. Exponential backoff is a retry mechanism where the time interval between attempts is doubled (e.g., first attempt = 1 second, second attempt = 2 seconds, third attempt = 4 seconds, fourth attempt = 8 seconds, etc.) or multiplied by some constant or exponent each time a failure occurs and the system is to retry the operation. It is used as an efficient mechanism to maintain retries when it is unknown how long the error or network downtime might last.

[0165] However, if there are remaining candidate systems 560, process 500 selects 570 the candidate system to perform failover. In one example, process 500 selects the candidate system with the best score (e.g., the lowest score in one example).

[0166] Process 500 determines whether the selected candidate system (572) is the current system. If the selected candidate system is the current system, process 500 ends (574). When the current system is available, the application is restarted on the current system. However, if the selected candidate system is not the current system, process 500 initiates a failover (576) on the selected candidate system. In one example, this includes initiating application deployment on the selected candidate system to deploy and execute the application on the selected candidate system. Deployment / execution on the selected candidate system includes, for example, taking steps to make the target application operational on the candidate system. This may include, for example: mounting a data volume that has been copied to the candidate system, providing data to the candidate system from a remote source, recreating system components and configuring them, etc.

[0167] An example of pseudocode used for the recovery process includes, for example:

[0168] # systems: includes all failover system targets and current system

[0169] systemScores = {}

[0170] for system in systems:

[0171] values ​​= getSystemValues(system) # values: dictionary of all scoredstatistics

[0172] systemScores[system] = scoreSystem(values)

[0173] systemBest = indexOf(min(systemScores))

[0174] if systemBest != currentSystem:

[0175] startFailover(systemBest)

[0176] # implicit else: stay on current system; do not fail over.

[0177] Examples of using the recovery process include:

[0178] Example A: Weights: Recovery Point Actual Data (RPA) = 50%; Past Recovery Time = 10%; Completed Application Recovery = 10%; Current Resource Availability = 15%; Idle CPU, Memory, Disk, and Network Latency are each 25% of Current Resource Availability.

[0179] System statistics for Site 1, Site 2, Site 3, Site 4

[0180] RPA 5 1 10 1.1

[0181] Past recovery time: 30 15 10 20

[0182] Completed application recovery: 50 100 25 75

[0183] Idle CPU cores: 12, 8, 16, 10

[0184] Free memory (GB) 8 8 128 4

[0185] Free disk space (GB / s): 8, 2, 16, 1

[0186] Network latency (ms): 43, 10000, 70, 10

[0187] Final score 23.4499 12.1038 26.0694 17.1858

[0188] In this example, Site 2 is the current site, and based on the final rating (where the best rating is the lowest rating compared to the others), the recommendation is to keep the application on the current site (e.g., Site 2). While this example includes some systematic statistics, additional, fewer, and / or other systematic statistics and / or other criteria can be considered. Furthermore, although some values ​​are used, others can be used. Many examples and variations are possible.

[0189] Example B: Weights: Recovery Point Actual Availability (RPA) = 50%; Past Recovery Time = 10%; Completed Application Recovery = 10%; Current Resource Availability = 15%; Idle CPU, Memory, I / O Bandwidth, and Network Latency are all 25% of the current resource availability.

[0190] System statistics for Site 1, Site 2, Site 3, Site 4

[0191] RPA 5 1 10 1.1

[0192] Past recovery time: 30 15 10 20

[0193] Completed application recovery: 50 100 25 100

[0194] Idle CPU cores: 12, 8, 16, 10

[0195] Free memory (GB) 8 8 128 4

[0196] Free disk space (GB / s): 8, 2, 16, 1

[0197] Network latency (ms): 43, 10000, 70, 10

[0198] Final score 23.4499 12.1038 26.0694 10.9358

[0199] In this example, site 2 is the current site, and based on the final score (where the best score is the lowest score compared to the others), a failover to site 4 is recommended. While some system statistics are included in this example, additional, fewer, and / or other system statistics and / or other criteria can be considered. Furthermore, although some values ​​are used, others can be used. Many examples and variations are possible.

[0200] Other examples, including additional, fewer, and / or other sites; additional, fewer, and / or other system statistics and / or other criteria; and / or other values, are possible. Many variations are possible.

[0201] In one or more aspects, a capability is provided to facilitate processing within a computing environment by providing intelligent selection of disaster recovery sites (candidate systems) for stateful applications (e.g., applications with state). The recovery process considers multiple objectives when determining which candidate system to select to execute the application based on selected events. These objectives include, for example, minimizing data loss and application downtime, and maintaining recovery and / or post-recovery performance. To facilitate achieving these objectives, one or more system statistics, one or more user-specified system statistics, and / or other criteria are considered. In one or more aspects, the recovery process collects, estimates, normalizes, and weights factors, and selects candidate systems to execute the application based on these factors.

[0202] In one or more respects, the recovery process is based on a deliberately transparent template. The template provides instructions on the information to be obtained and considered (e.g., objectives, system statistics, user-specified system statistics, factors, and / or other criteria), as well as the interface for obtaining that information. It clearly outlines the information to be obtained and how to use it.

[0203] In one or more respects, applications are stateful because there is application state (data) that can be lost. As an example, application state can be disaster-protected using, for example, backup / disaster recovery, where the state is replicated to one or more pre-selected intermediate locations after the event and restored to a post-event selected application host system (the selected system); and / or disaster recovery, where the state is replicated to a pre-selected system capable of hosting the event. One or more aspects of the recovery process described herein are unrelated to these disaster protection techniques.

[0204] In one or more aspects, multiple possible failure scenarios are provided across multiple sites (candidate systems), along with conditional logic for selecting the best available option from the multiple sites (candidate systems).

[0205] One or more aspects of this disclosure relate to computer technology and facilitate processing within a computer, thereby improving its performance. For example, processing within a computing environment is improved by providing ease of recovery capabilities, which reduce application performance degradation and / or data loss and improve recovery time. In one or more aspects, one or more applications on a specific system are recovered based on a selected event (e.g., a natural disaster) affecting that specific system of the computing environment. Recovery includes automatically selecting a candidate system from a plurality of candidate systems based on the selected event, in which the one or more applications are executed. The selected candidate system can be the same system or another candidate system. If the same system is selected, one or more actions are performed to enable the application to run on that candidate system. For example, restarting, repairing and / or physically moving the candidate system, mounting volumes, performing backups and / or taking any other actions to make the candidate system usable. Based on the candidate system being physically ready to run one or more applications, one or more applications are launched and executed on the candidate system. By selecting the same system as the candidate system, data loss is eliminated (or significantly minimized), saving processing time and system resources. For example, processing time and system resources are saved by avoiding redundant processing to recreate data and / or by avoiding tasks that could be performed to deploy one or more applications on different candidate systems.

[0206] If a candidate system different from the specific system affected by the selected event is selected, one or more actions are performed to enable the deployment and execution of one or more applications on the selected candidate system. For example, one or more data volumes can be installed, data and / or components / modules can be provided to the selected candidate system and installed, and / or one or more applications can be installed, etc. Subsequently, one or more applications are executed on the selected candidate system. The selected candidate system can improve processing within the computer by executing one or more applications. Performance within the computing environment is improved by quickly recovering and executing one or more affected applications. Based on selecting a candidate system that minimizes data loss, performance within the computing environment is improved because repetitive processing to recreate data is minimized, saving processing time and resources within the computer. Processing within the processor, computer system, and / or computing environment is improved.

[0207] In one or more respects, technological advancements such as cloud technology, which enable certain resources such as storage to be highly available without being tied to a specific location and / or to be highly available in the event of varying degrees of data loss depending on location, demonstrate the usefulness of such recovery solutions that dynamically select candidate systems for failover based on multiple criteria or metrics.

[0208] In one or more aspects, the multi-criteria problem has been addressed, which is increasingly beneficial for recovery points in multiple regions, including the cloud.

[0209] In one or more aspects, a system is recommended to recover applications during a disaster by considering, for example, the actual recovery point, predicted recovery time, and historical data of the recovery system. Even if the Recovery Point Objective (RPO) is met, data loss is limited by considering the actual Recovery Point Assessment (RPA) range. For example, the RPO may be 15 minutes, but the RPA for system A is 3 minutes, and for system B it is 14 minutes. Furthermore, in one or more aspects, recovery performance is estimated based on historical measurements. Additionally, in one or more aspects, a disaster recovery plan is executed by selecting a chosen fault recovery system (candidate system) that meets or exceeds each objective, weighted relative to other factors and estimated by a function.

[0210] In one or more respects, by taking into account, for example, the lifetime of each data copy and the current state of the system to which it is attached, data loss and application downtime during disasters are reduced.

[0211] One or more aspects are optimized for certain objectives, such as minimizing application downtime, minimizing data loss, and maintaining application performance after failover. In one or more aspects, the user can select certain information to consider during failover. For example, the user can select one or more system statistics and / or user-defined information. User-defined information includes, for example, statistics of one or more user identifiers, such as one or more system statistics and / or other statistics (e.g., time of day, etc.), and the user expects that the relative importance of the statistics, along with the user identifier, has been taken into account. A standardized method is provided to score the selected information and use these scores to select candidate systems. Furthermore, different values ​​can be compared between systems for fair comparison, and a weighting scheme that can be customized by the user according to their needs is provided.

[0212] In one or more aspects, capabilities are not limited to a specific configuration, and additional criteria are considered when making failover recommendations; for example, optimizations for minimizing data loss, minimizing recovery time, etc. These one or more aspects can be generalized across different users with different needs.

[0213] In one or more aspects, one or more criteria are considered that can be used to make decisions, including, for example, minimizing application downtime and maintaining application performance continuity after workload migration. Furthermore, a weighting mechanism for failover attributes is provided, allowing users to emphasize or de-emphasize characteristics desired by a particular application, and additional criteria are defined for evaluating failover decisions. In one or more aspects, criteria are continuously evaluated, and current and historical performance levels are included when making decisions.

[0214] In one or more aspects, multiple failover options are provided, and in the event of a disaster, the best failover target recovery system (candidate system) is selected.

[0215] Other aspects, variations, and / or embodiments are possible.

[0216] In addition to the above, one or more aspects such as providing, supplying, deploying, managing, and servicing customer environment management can be provided by a service provider. For example, a service provider can create, maintain, support, etc., the computer code and / or computer infrastructure that performs one or more aspects for one or more customers. In return, the service provider can receive payments from customers, for example, under subscription and / or fee agreements. Alternatively or alternatively, the service provider can receive payments from selling advertising content to one or more third parties.

[0217] In one aspect, an application can be deployed to execute one or more embodiments. As an example, application deployment includes providing computer infrastructure operable to execute one or more embodiments.

[0218] On the other hand, computing infrastructure can be deployed, including integrating computer-readable code into a computing system, wherein the code combined with the computing system is capable of executing one or more embodiments.

[0219] On the other hand, a process for integrating computing infrastructure can be provided, including integrating computer-readable code into a computer system. The computer system includes a computer-readable medium, wherein the computer medium includes one or more embodiments. The code integrated with the computer system is capable of executing one or more embodiments.

[0220] While various embodiments have been described above, these are merely examples. For instance, other techniques may be used to select candidate systems, perform recovery, and / or perform one or more other aspects of this disclosure. Many variations are possible.

[0221] This document describes various aspects and embodiments. Furthermore, many variations are possible without departing from the scope of the claims. It should be noted that, unless otherwise inconsistent, each aspect or feature described and / or claimed herein, and its variations, may be combined with any other aspect or feature.

[0222] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used herein, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the terms “comprising” and / or “including” as used in this specification specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0223] If present, all means or steps plus functional elements in the following claims are intended to include corresponding structures, materials, actions, and equivalents for performing functions in combination with other claimed elements of the particular claim. Descriptions of one or more embodiments have been presented for purposes of illustration and description, but such description is not intended to be exhaustive or limited to the forms disclosed. Many modifications and variations will be apparent to those skilled in the art. The embodiments were chosen and described in order to best explain various aspects and practical applications, and to enable others skilled in the art to understand the various embodiments with various modifications suitable for the particular intended use.

Claims

1. A computer program product for facilitating processing within a computing environment, the computer program product comprising: At least one computer-readable storage medium and program instructions commonly stored on the at least one computer-readable storage medium, wherein the commonly stored program instructions include: Program instructions for obtaining user-defined information related to the recovery of the application based on the selected event; Program instructions for obtaining data related to a set of candidate systems that can be used to deploy the application, the data including one or more system statistics for the set of candidate systems, the one or more system statistics including at least one or more recovery point actual data and historical data for the set of candidate systems; Program instructions for performing a scoring of the candidate system set based on the user-defined information and the obtained data; Program instructions for selecting candidate systems from the candidate system set based on the score; and Program instructions used to initiate the deployment of the application on the selected candidate system.

2. The computer program product according to claim 1, wherein, The user-defined information includes the relative importance of at least one system statistic from one or more candidate systems in the candidate system set to the user.

3. The computer program product according to claim 1 or claim 2, wherein, The user-defined information includes statistics on one or more user identifiers.

4. The computer program product according to claim 3, wherein, The statistics for identifying one or more users include: an indication of a valuation function for estimating at least one of the system statistics of one or more candidate systems in the candidate system set.

5. The computer program product according to any one of the preceding claims, wherein, The one or more system statistics also include: one or more resource availability measures for the system resource set of at least one candidate system in the candidate system set.

6. The computer program product according to claim 5, wherein, The system resource set includes memory, one or more central processing units, input / output bandwidth, and latency.

7. The computer program product according to any one of the preceding claims, wherein, The statistics for one or more systems also include the degree of complete recovery of the application.

8. The computer program product according to any one of the preceding claims, wherein, The program instructions for obtaining the data include: program instructions for obtaining the data based on the occurrence of the selected event.

9. The computer program product according to any one of the preceding claims, wherein, The candidate system set includes the current system on which the application is deployed, as well as one or more other candidate systems.

10. The computer program product according to any one of the preceding claims, wherein, The application is a stateful application, and the application data of the application is protected using one or more data protection technologies, wherein the selection of the candidate system is independent of the one or more data protection technologies.

11. The computer program product according to any one of the preceding claims, wherein, The score is based on selected data related to application downtime, current resource availability of the candidate system set, time taken to restore the application, post-recovery performance, and estimated post-recovery resource availability.

12. The computer program product according to any one of the preceding claims, wherein, The program instructions for performing the scoring include: program instructions for performing weighting of at least one system statistic among the one or more system statistics based on the user-defined information to obtain one or more weighted values, and wherein the program instructions for performing the scoring include: program instructions for using the one or more weighted values ​​in the scoring.

13. The computer program product according to any one of the preceding claims, wherein, The historical data includes one or more historical recovery time statistics for one or more candidate systems in the candidate system set.

14. The computer program product according to any one of the preceding claims, wherein, The one or more recovery point actual data include at least one recovery point actual data for each candidate system in the candidate system set.

15. The computer program product according to any one of the preceding claims, wherein, The selected event is disaster recovery from a natural disaster.

16. A computer system for facilitating processing within a computing environment, the computer system comprising: Memory; as well as At least one device coupled to the memory, wherein the computer system is configured to perform a method comprising: Based on the selected event, obtain user-defined information related to the application's recovery; Obtain data related to a set of candidate systems that can be used to deploy the application, the data including one or more system statistics for the set of candidate systems, the one or more system statistics including at least one or more recovery point actual data and historical data for the set of candidate systems; The candidate system set is scored based on the user-defined information and the obtained data. Candidate systems are selected from the candidate system set based on the scores; and Initiate the deployment of the application on the selected candidate system.

17. The computer system according to claim 16, wherein, The user-defined information includes the relative importance of at least one system statistic from one or more candidate systems in the candidate system set to the user.

18. The computer system according to claim 16 or 17, wherein, The score is based on selected data related to application downtime, current resource availability of the candidate system set, time taken to restore the application, post-recovery performance, and estimated post-recovery resource availability.

19. The computer system according to any one of claims 16 to 18, wherein, Performing the scoring includes: weighting at least one of the one or more system statistics based on the user-defined information to obtain one or more weighted values, and wherein performing the scoring includes: using the one or more weighted values ​​in the scoring.

20. The computer system according to any one of claims 16 to 19, wherein, The application is a stateful application, and the application data of the application is protected using one or more data protection technologies, wherein the selection of the candidate system is independent of the one or more data protection technologies.

21. A computer-implemented method for facilitating processing within a computing environment, the computer-implemented method comprising: Based on the selected event, obtain user-defined information related to the application's recovery; Obtain data related to a set of candidate systems that can be used to deploy the application, the data including one or more system statistics for the set of candidate systems, the one or more system statistics including at least one or more recovery point actual data and historical data for the set of candidate systems; The candidate system set is scored based on the user-defined information and the obtained data. Candidate systems are selected from the candidate system set based on the scores; as well as Initiate the deployment of the application on the selected candidate system.

22. The computer-implemented method according to claim 21, wherein, The user-defined information includes the relative importance of at least one system statistic from one or more candidate systems in the candidate system set to the user.

23. The computer-implemented method according to claim 21 or 22, wherein, The application is a stateful application, and the application data of the application is protected using one or more data protection technologies, wherein the selection of the candidate system is independent of the one or more data protection technologies.

24. The computer-implemented method according to any one of claims 21 to 23, wherein, The score is based on selected data related to application downtime, current resource availability of the candidate system set, time taken to restore the application, post-recovery performance, and estimated post-recovery resource availability.

25. The computer-implemented method according to any one of claims 21 to 24, wherein, Performing the scoring includes: weighting at least one of the one or more system statistics based on the user-defined information to obtain one or more weighted values, and wherein performing the scoring includes: using the one or more weighted values ​​in the scoring.