Mainframe Failure Prediction via Tensor Time Series Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for monitoring mainframe computer system failures are reactive, leading to delayed detection and excessive downtime, as they rely on rough estimates of failure times that do not accurately predict hardware-related failures, resulting in inefficient servicing and potential increased downtime.
Innovation Solution
A method is introduced to structure raw data into dynamic or fixed-sized tensor sets, which are processed in a time series manner to teach machine learning algorithms to identify patterns leading to failures, allowing for proactive detection and prediction of mainframe computer system failures.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional failure monitoring methods are used, then the system can detect failures after they occur, but the response time is delayed and downtime is excessive
Solution Approach 1:
The system performs preliminary analysis of operational data to identify patterns and anomalies that precede failures. By analyzing data trends before actual failure occurs, the system can predict potential failures and schedule maintenance proactively, thereby reducing downtime while maintaining accurate failure detection.
Solution Approach 2:
The system implements continuous feedback loops where operational data is constantly monitored, analyzed, and used to update failure predictions. This feedback mechanism allows the system to learn from past failures and improve its prediction accuracy over time, enabling earlier detection and response to potential failures.
2Productivity
If rough estimates of failure times are used, then servicing can be scheduled, but unnecessary servicing occurs disrupting normal operations
Solution Approach 1:
The system replaces rough mechanical estimation methods with advanced data analysis and pattern recognition algorithms. By substituting simple time-based estimates with sophisticated operational data analysis, the system achieves both accurate failure prediction and minimal disruption to normal operations.
Solution Approach 2:
The system changes the parameters used for failure prediction from simple time-based rough estimates to multiple operational parameters including performance metrics, error rates, and environmental conditions. This multi-parameter approach enables more accurate predictions that reduce unnecessary servicing while maintaining operational efficiency.
3Reliability
If data is structured for machine learning analysis, then proactive failure detection is enabled, but data processing complexity increases
Solution Approach 1:
The system segments operational data into distinct features and dimensions that are relevant for failure prediction. By dividing raw data into meaningful segments such as performance metrics, error patterns, and operational conditions, the system enables effective machine learning analysis while managing data processing complexity through structured organization.
Data Source
AI summary
A system and method for generating a data set structured for recognition of time series data by a machine learning computer are provided. The method includes acquiring time series data, generating tensor units based on the time series data, and identifying a target tensor unit including a time of failure of a mainframe computer system. The method further includes generating tensor sets, in which at least one tensor set includes the target tensor unit. The generated tensor sets are then migrated to a machine learning computer for generating or updating of a computer model based on the time series data, the computer model recognizing a data pattern preceding the time of failure of the mainframe computer system. The computer model is then applied to data in a production environment for identifying a production data pattern corresponding to a data pattern recognized in the tensor sets.


