Application Availability Optimization via Failure Profile Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Distributed applications face challenges in determining an optimal configuration that balances maximum protection with minimum costs, as IT infrastructure components can fail, impacting the entire application and varying in failure rates, leading to different levels of failure classification.
Innovation Solution
A method and system that compute the actual application impact and failure profile based on the number of failing IT infrastructure components, determining a factor in the likelihood of failure, and analyzing scenarios to optimize application architecture for maximum protection with minimal costs by assessing different configurations such as redundancy and component placement.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If redundancy is added to IT infrastructure components to improve application availability, then reliability is improved, but cost increases
Solution Approach 1:
The system performs preliminary failure analysis and simulation before implementing redundancy measures. By computing failure profiles and identifying critical components in advance, the system determines the optimal level of redundancy needed, avoiding unnecessary duplication of non-critical components while ensuring adequate protection for critical ones.
Solution Approach 2:
The system applies different levels of redundancy to different IT infrastructure components based on their individual failure profiles and impact on application availability. Critical components with high failure impact receive higher redundancy, while less critical components receive minimal or no redundancy, optimizing the balance between reliability and cost.
2Reliability
If more IT infrastructure components are used to improve application protection, then reliability is improved, but device complexity increases
Solution Approach 1:
The system performs preliminary analysis to identify the minimum necessary components for achieving target availability levels. By simulating failures and computing impact profiles in advance, the system determines the simplest architecture that meets reliability requirements, avoiding unnecessary complexity.
Solution Approach 2:
The system segments the IT infrastructure into critical and non-critical components based on failure impact analysis. This segmentation allows for simplified management and configuration, where redundancy is applied only to segmented critical components, reducing overall system complexity while maintaining necessary protection.
3Measurement precision
If failure analysis is performed on all components, then measurement precision is improved, but loss of time increases
Solution Approach 1:
The system focuses failure analysis efforts on specific critical components rather than uniformly analyzing all components. By identifying and prioritizing components with the highest impact on application availability, the system achieves high measurement precision for critical areas while minimizing time spent on less important components.
Solution Approach 2:
The system performs partial failure analysis on the most critical components rather than exhaustive analysis of all components. This partial action approach achieves sufficient measurement precision for decision-making while significantly reducing the time required for the analysis process.
Data Source
AI summary
An approach to an optimal application configuration. The approach includes a method that includes computing, by at least one computing device, an actual application impact based on an “N” number of failing information technology (IT) infrastructure components within an application architecture. The method includes determining, by the at least one computing device, a factor in likelihood of failure of the “N” number of IT infrastructure components. The method includes determining, by the at least one computing device, a failure profile for the application architecture based on the actual application impact and the factor in likelihood of failure.


