Off-Cluster Logset Health Evaluation for Targeted Server Auto-Remediation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for resolving operational issues in data servers within a server cluster are manual, time-consuming, and often result in delayed resolution, causing disruption to digital operations and requiring extensive customer engagement for software service packs, which are generated and distributed belatedly.
Innovation Solution
A system that automatically identifies operational issues in data servers, compares them with historical issues, and implements corrective actions without affecting other servers in the cluster, using artificial intelligence and machine learning to generate and distribute timely resolutions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual methods are used to resolve operational issues in data servers, then customer engagement and software service packs are required, but the resolution is delayed and causes disruption to digital operations
Solution Approach 1:
The system enables servers to automatically diagnose and resolve their own operational issues by comparing logsets against a knowledge base of historical issues and resolutions, eliminating the need for manual customer engagement and accelerating resolution while minimizing downtime
Solution Approach 2:
The system pre-populates a knowledge base with historical operational issues and their resolutions before they occur, enabling rapid automated matching and resolution when similar issues arise, thus reducing resolution time and preventing service disruption
2Productivity
If software service packs are generated and distributed manually, then customer engagement is required, but the process is time-consuming and results in belated resolution
Solution Approach 1:
The system replaces the manual mechanical process of generating and distributing software service packs with an automated electronic system that uses AI/ML to match operational issues against a knowledge base and automatically generate and apply resolutions, dramatically increasing resolution speed and eliminating delays
3Productivity
If automated actions are implemented on a data server, then issue resolution is accelerated, but other servers in the cluster may be affected
Solution Approach 1:
The system segments the server cluster into individual server units, enabling automated diagnostic and remediation actions to be applied exclusively to the affected server through targeted logset analysis and action generation, preventing propagation of issues to other servers and maintaining cluster stability
Data Source
AI summary
Various systems and methods are presented herein regarding identifying an operational issue are a data server, automatically identifying/implementing an action to fix the operational issue. The data server can be co-located with a collection of data servers in a server cluster. The action can be configured to be specifically implemented at the data server without affecting an operational status of the other data servers in the collection of data servers. The action can be a server reboot/reset instruction, terminate operation of an application, and suchlike. The operational issue can be compared with a prior operational issue having an associated action, wherein the associated action can be utilized as the action to fix the operational issue at the data server. Over time, respective actions implemented at the one or more data servers in the server cluster can be compiled from which a software service pack can be subsequently compiled and distributed.


