Data Protection Advisor for Human Configuration Error Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In large data centers, determining the cause of errors or failures is challenging, particularly when they are related to human configuration changes during periodic events like backups or data replication, as existing methods lack efficiency in isolating and resolving such errors quickly.
Innovation Solution
A data protection advisor application automatically detects errors by analyzing configuration, status, and event data, identifying configuration changes made by human users, and provides feedback with suggested resolutions, enabling quick identification and remediation of error causes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual error detection methods are used in large data centers, then operators can identify errors, but the process is time-consuming and inefficient for isolating human configuration errors during periodic events
Solution Approach 1:
The system performs self-diagnosis by automatically collecting configuration data, event data, and status data, then analyzing them to identify human configuration errors without requiring manual operator intervention for each error detection case
Solution Approach 2:
Manual mechanical error detection processes are replaced with an automated electronic system that collects data from multiple sources, processes it through analysis algorithms, and generates error identification results automatically
2Measurement precision
If comprehensive data collection is performed to accurately identify error causes, then detection accuracy improves, but system complexity increases
Solution Approach 1:
The data collection system is segmented into distinct modules: a configuration data collector that gathers configuration information, an event data collector that captures periodic event data, and a status data collector that monitors system status, with each module independently collecting specific data types
Solution Approach 2:
An analysis module acts as an intermediary that receives data from multiple collection modules, processes the combined information, and generates error identification results, thereby managing the complexity of integrating comprehensive data from diverse sources
Data Source
AI summary
Detection of a cause of an error occurring in a data center is provided. A time period between a failure event in the data center and a previous successful event is determined. A set of devices involved in the failure event is generated using a data protection advisor (DPA). Data collected for the set of devices during the determined time period is scanned to detect at least one configuration change made during the determined time period. Based on a result of the scanned collected data, the at least one configuration change is displayed as a potential root cause of the error.


