FRU Non-Volatile Storage for Fault Analysis Data Logging
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In electronic systems, hardware failures often go undocumented when devices are replaced in the field, leading to inefficient troubleshooting and potential reintegration of defective parts, as service personnel lack time to gather failure data, and existing fault management tools struggle to accurately predict device failures post-deployment.
Innovation Solution
Implementing a fault analysis system that writes failure information to non-volatile storage on field replaceable units (FRUs), allowing the system to identify faulty components and document symptoms and reasons for replacement, thereby providing valuable data for manufacturers and service organizations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If service personnel manually gather and document failure data when replacing hardware devices, then fault information accuracy is improved, but service time and system downtime increase
Solution Approach 1:
The system performs preliminary action by automatically capturing and storing fault information in non-volatile memory of the FRU at the moment of failure detection, before service personnel arrive. This eliminates the need for manual data gathering during replacement, resolving the contradiction between information accuracy and service time.
Solution Approach 2:
The FRU performs self-service by automatically documenting its own failure information and storing it locally. The device self-diagnostics and self-documented fault conditions, symptoms, and replacement reasons without requiring external intervention, thus improving accuracy while minimizing service time.
2Productivity
If fault information is not documented on replaced devices, then service process speed is improved, but troubleshooting efficiency and failure model accuracy deteriorate
Solution Approach 1:
The system captures and stores fault information in advance at the moment of failure, before the device is removed from service. This preliminary documentation ensures that complete fault data is preserved without requiring additional time during the replacement process, thus maintaining service speed while preventing information loss.
Solution Approach 2:
The non-volatile memory in the FRU acts as an intermediary that automatically stores and preserves fault information. This intermediary component captures diagnostic data, symptoms, and replacement reasons, ensuring information is retained for later analysis without affecting the speed of the service process.
3Adaptability or versatility
If existing fault management tools are used to predict device failures, then failure prediction capability is improved, but accuracy in predicting actual failures post-deployment deteriorates
Solution Approach 1:
The system implements feedback by collecting actual failure data from deployed devices through automatic documentation and using this real-world data to refine and improve failure prediction models. The documented fault information from actual failures feeds back into updating the prediction algorithms, increasing their accuracy for future predictions.
Solution Approach 2:
The system changes parameters by using actual field failure data to adjust and refine failure prediction model parameters. The documented symptoms, conditions, and failure patterns from real deployments are used to modify prediction thresholds and parameters, improving the models' accuracy for predicting actual failures in similar conditions.
Data Source
AI summary
A system and method for recording fault information in an electronic system are disclosed herein. A system includes fault analysis logic and a plurality of field replaceable units (“FRUs”). The fault analysis is configured to analyze system error information, and identify at least one of the FRUs in the system to be a possible cause of a detected fault based on the analysis. Each FRU includes writeable non-volatile storage including storage locations reserved to store information including a result of the analysis. The result of the analysis indicates a reason that the FRU storing the information was determined, by the fault analysis logic, to be a possible cause of the fault.


