Incident-Responsive Snapshot Generation for Transient Fault Diagnosis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional ITSM systems face challenges in diagnosing transient technical issues in computing devices due to incomplete and unreliable information, as the state of the user device changes between the submission of a support ticket and the IT support provider's intervention, masking important parameters necessary for issue diagnosis.
Innovation Solution
The implementation of incident-responsive computing system snapshot generation, which captures and compares computing parameters at the time of the technical issue, including transient parameters, to identify the cause of the issue and enable remote mitigation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional ITSM systems wait for support ticket submission before diagnosing issues, then support providers can address user-reported problems, but the device state changes between ticket submission and diagnosis, causing loss of transient parameters needed for accurate diagnosis
Solution Approach 1:
The system proactively captures computing parameters and generates snapshots of device state before the support provider needs to diagnose the issue. This preliminary data collection ensures transient parameters are preserved, allowing accurate diagnosis when the support provider reviews the captured data later.
Solution Approach 2:
The system creates copies of the device state through computing parameter snapshots. These snapshots capture the complete device state at specific moments, allowing support providers to analyze replicated data without needing to access the actual device during the diagnostic process, thus preserving transient information.
2Loss of information
If the system captures computing parameters continuously to ensure complete diagnostic information, then accurate diagnosis of transient issues is achieved, but system complexity and resource consumption increase
Solution Approach 1:
Instead of continuous monitoring, the system captures computing parameters periodically or at predetermined intervals. This approach preserves transient information while significantly reducing system complexity and resource consumption compared to continuous capture, as data is collected only at specific moments rather than constantly.
Solution Approach 2:
The system selectively captures specific computing parameters relevant to diagnosis rather than all possible device data. This targeted approach focuses resources on collecting diagnostically valuable information while minimizing overall system complexity and resource usage.
3Measurement precision
If the system captures computing parameters immediately responsive to incident signals, then transient technical states are accurately recorded for diagnosis, but response time and processing overhead increase
Solution Approach 1:
The system captures computing parameters immediately when incident signals are received, before the device state can change. This timely capture preserves the exact transient state associated with the incident, enabling accurate diagnosis without significant time loss, as the data is collected at the critical moment rather than later.
Data Source
AI summary
A method of remote device diagnosis and mitigation includes receiving a signal indicative of an intermittent technical state of a first device. Immediately responsive thereto, the method includes interrogating the first device for parameters. The method includes interrogating the first device for the parameters at a third time outside receipt of the signal. The parameters include a transient parameter present at a first time of the intermittent technical state and not present a second time following the first time. The method includes recording the parameters from the first time in a first data file and the parameters for the third time in an additional data file. The first data file is compared with the additional data file to identify a difference in a parameter indicative of a cause of the intermittent technical state. The method includes remotely implementing a change on the first device to mitigate the cause.


