Dynamic Workload Allocation via Real-Time Resiliency Scoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current workload allocation methods in networked computing environments are inadequate for dynamic failover and resilience management, as they rely on static allocation and lack the ability to assess the resilience of potential failover servers, leading to potential failures and inefficient workload distribution.
Innovation Solution
The solution involves obtaining risk factor data for components in a global data center, calculating resiliency scores based on applicable risk factors, aggregating scores for component groups, and selecting datacenters for failover protection based on group resiliency scores to improve enterprise resilience and application availability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If static workload allocation is used, then device complexity is reduced, but reliability deteriorates due to inability to dynamically respond to failures
Solution Approach 1:
The patent implements dynamic workload allocation by transitioning from static pre-configured failover to real-time decision-making based on resiliency scoring. The system continuously evaluates multiple datacenters' resilience status and dynamically redirects workloads to the most suitable available datacenter, ensuring optimal application availability without requiring complex manual intervention.
Solution Approach 2:
The system employs feedback mechanisms by calculating and monitoring resiliency scores for each datacenter based on real-time risk factor data. This feedback loop enables the workload allocation system to assess current datacenter health, compare alternatives, and make informed dynamic allocation decisions, thereby improving reliability through data-driven automation.
2Reliability
If comprehensive risk factor assessment is implemented, then reliability is improved, but device complexity increases due to multiple scoring components
Solution Approach 1:
The resiliency assessment system is segmented into distinct risk factor categories (e.g., infrastructure risk, application risk, geographic risk) that can be independently evaluated and scored. This segmentation allows the complex assessment to be broken down into manageable components, each contributing to the overall resiliency score while maintaining system tractability.
Solution Approach 2:
The system utilizes parameter changes by incorporating multiple variable risk factors that can be adjusted and weighted based on specific organizational needs and threat landscapes. This flexibility allows the scoring system to adapt to different scenarios and priorities without requiring complete system redesign, balancing comprehensiveness with manageability.
3Reliability
If dynamic failover selection is used, then reliability is improved, but loss of time increases due to real-time evaluation requirements
Solution Approach 1:
The system performs preliminary actions by pre-establishing resiliency scoring frameworks, risk factor methodologies, and datacollection mechanisms before failures occur. This preparation enables rapid real-time evaluation during actual failover events, as the foundational assessment infrastructure is already in place and does not need to be constructed during critical transition moments.
Data Source
AI summary
An approach is provided for providing disaster recovery, active/active and active/standby workload allocation in a networked computing environment according to aspects of the present invention. Risk factor data is obtained for each component of a plurality of components in a global data center. A resiliency score is calculated for each component of the plurality of components based on a set of risk factors that are applicable to the component gathered from the risk factor data. A set of group resiliency scores are computed by aggregating the resiliency scores for each of the plurality of components included in a component group. In response to a determination that an application's performance can be improved, a datacenter is selected for failover protection based on a group resiliency score corresponding to the datacenter. Moreover, the overall enterprise resiliency score can be improved by moving an application between sites in the enterprise.


