Node Subset Correction for Web Service Performance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current web service monitoring systems fail to effectively address performance issues in distributed web applications without causing significant service disruptions, leading to losses for service providers.
Innovation Solution
A system that monitors business transactions and applies corrective actions to a subset of nodes experiencing performance issues, using health rules and policies to determine the necessary actions, allowing for targeted problem-solving and minimizing service downtime.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If corrective actions are applied to all nodes experiencing performance issues, then performance problems are fully addressed, but service disruption increases significantly
Solution Approach 1:
The system segments the set of affected nodes into multiple subsets and applies corrective actions to one subset at a time rather than all nodes simultaneously. This allows performance issues to be addressed systematically while maintaining service availability on other node subsets, thereby reducing overall service disruption and downtime.
Solution Approach 2:
The system applies corrective actions partially to only a subset of affected nodes rather than all nodes. By correcting performance issues on a portion of nodes first, the system can restore service capacity progressively while avoiding the complete service disruption that would result from applying corrections to all nodes at once.
2Reliability
If administrators manually take applications offline to correct problems, then performance issues are resolved, but service availability is significantly reduced
Solution Approach 1:
The system implements automated monitoring and self-service corrective actions that detect performance issues and apply corrections without requiring administrator intervention to take applications offline. The automated system can apply corrective actions to node subsets while maintaining service availability, eliminating the need for manual service disruption.
Solution Approach 2:
The system performs preliminary monitoring and detection of performance issues before they escalate, and has corrective actions pre-configured and ready to apply. This allows the system to address performance problems proactively while maintaining service availability, rather than reacting by taking applications offline after problems occur.
3Reliability
If multiple applications are taken offline simultaneously, then all performance problems are addressed, but business loss and customer loyalty are significantly impacted
Solution Approach 1:
The system segments the correction process into multiple phases, addressing performance issues in different node subsets at different times. This staggered approach ensures that some applications remain available to serve customers while others are being corrected, thereby minimizing business loss and maintaining customer loyalty while still resolving all performance problems.
Solution Approach 2:
The system maintains continuous service availability by ensuring that while some node subsets are being corrected, other node subsets continue to provide service. This continuity of useful action prevents complete service interruption, minimizing business loss and customer impact while performance corrections are being applied.
Data Source
AI summary
Business transactions and the nodes processing the transactions are monitored and actions are applied to one or more nodes when a performance issue is detected. A performance issue may relate to a metric associated with a transaction or node that processes the transaction. If a performance metric determined from data captured by monitoring does not satisfy a health rule, the policy determines which action should be performed to correct the performance of the node. When a problem is detected for multiple nodes, the present technology may address a subset of the multiple nodes rather than apply an action to each node experiencing the problem. When a solution is found to correct the problem with the subset of nodes, the solution may be applied to the other nodes experiencing the same problem.


