Data Estimation and Forecasting Matrix Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for predicting and completing datasets in large equipment systems are limited by the presence of missing data points, which hinder accurate predictive modeling and future data forecasting without prior knowledge of equipment behavior.
Innovation Solution
A Data Estimation and Forecasting (DEF) computing device that uses principal component analysis, probabilistic principal component analysis, and other techniques to estimate missing data points and forecast future values by arranging data into matrices, generating sample matrices, applying normalization and scaling, and generating non-null values for missing variables.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If known methods are used to estimate missing data, then some data points can be predicted, but specific variables cannot be estimated in isolation and massive computing capacity is required
Solution Approach 1:
The patent segments the dataset into multiple matrices (primary matrix, sample matrix, augmented matrix) and processes them separately through distinct computational steps including parsing, sampling, augmentation, normalization, and scaling. This segmentation allows the system to handle large datasets without requiring massive computing capacity by breaking down the estimation problem into manageable portions that can be processed iteratively.
Solution Approach 2:
The system performs preliminary actions by generating sample matrices from complete rows and creating augmented matrices with synthetic data patterns before the actual estimation process. These pre-computed structures establish statistical relationships and variable correlations in advance, enabling accurate estimation of specific variables in isolation without requiring massive real-time computing resources.
2Reliability
If traditional forecasting methods are used, then future data can be predicted, but at least some a priori knowledge of past equipment behavior is required
Solution Approach 1:
The patent introduces an intermediary mechanism - the augmented matrix with synthetic data patterns - that mediates between existing data and future predictions. This intermediary structure captures statistical relationships and equipment behavior patterns without requiring explicit a priori knowledge, allowing the system to forecast future data by leveraging relationships learned from the augmented data structure rather than requiring pre-programmed domain knowledge.
Solution Approach 2:
The system applies parameter changes through normalization and scaling operations that transform the augmented matrix into a format suitable for forecasting. By dynamically adjusting statistical parameters during the processing pipeline, the system adapts to different equipment behaviors and data distributions without requiring pre-configured knowledge, enabling reliable forecasting across diverse scenarios.
3Measurement precision
If complete operational history is maintained, then predictive modeling accuracy improves, but data completeness is seldom achieved in practice
Solution Approach 1:
The patent converts the harmful effect of missing data into a benefit by using the patterns and relationships present in the available data to generate synthetic augmented rows that represent the missing information. Rather than treating incomplete datasets as a problem to be avoided, the system leverages the structure of available data to create meaningful estimations, transforming data loss into an opportunity for pattern-based inference and accurate predictive modeling.
Data Source
AI summary
A system for estimating data in large datasets for an equipment system is provided. The system includes a data estimation and forecasting (DEF) computing device. The DEF computing device arranges a dataset in a primary matrix and parses rows of the primary matrix and generates a sample matrix by selecting primary matrix rows having non-null values for each variable. The DEF computing device adds to the sample matrix rows that include non-null values for each variable except one. The DEF computing device generates normalized values for this augmented matrix, applies several techniques including probabilistic principal component analysis (PPCA) and Markov processes, and scales the augmented matrix to normalized values. The DEF computing device generates non-null values for the variable, scales the augmented matrix back to the sample matrix, and generates a forecast for the equipment system, directing a user to update logistics processes for the equipment system.


