Time Series Prediction Using Space Partitioning Data Structures
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current prediction techniques for data streams become computationally expensive and slow with large volumes of data, often resulting in delayed or ineffective predictions, especially in systems with high activity levels, such as database systems experiencing peak loads.
Innovation Solution
The proposed solution involves a prediction system that uses polyline simplification to reduce sample subsets, converts data to angular coordinates, and stores them in a space partitioning data structure like a k-dimensional tree, allowing for faster prediction times by identifying nearest neighbors and generating predicted samples efficiently.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If current prediction techniques are used on large volumes of data, then prediction accuracy may be maintained, but processing time becomes excessively long and predictions become delayed
Solution Approach 1:
The data stream is divided into fixed-size sliding windows that partition historical data into manageable subsets. Each window contains a fixed number of recent data points, allowing the system to process only relevant recent history rather than the entire data stream. This segmentation enables efficient computation while maintaining prediction accuracy by focusing on the most relevant recent patterns.
Solution Approach 2:
The patent extracts only the essential features from historical data by using fixed-size windows that capture recent patterns without storing entire historical sequences. By extracting and processing only the necessary recent data points within each window, the system avoids the computational burden of processing all historical data while maintaining the ability to make accurate predictions.
2Reliability
If regression techniques are applied to analyze large data streams, then prediction capability is maintained, but computational complexity increases significantly
Solution Approach 1:
The data is segmented into fixed-size sliding windows, which limits the amount of historical data processed for each prediction. This segmentation reduces computational complexity by ensuring that regression analysis is performed only on a manageable subset of recent data points rather than the entire historical data stream, while still maintaining prediction capability.
Solution Approach 2:
The system performs preliminary data organization by maintaining pre-structured sliding windows of historical data in memory. This preliminary organization allows the regression algorithm to quickly access and process only the necessary recent data points without having to retrieve and organize large volumes of historical data during prediction, thereby reducing computational complexity.
Data Source
AI summary
Techniques are disclosed for a computer system to predict a next sample for a data stream that specifies data values of one or more variables. A current subset of data values and previous subsets of data values is determined, and polyline simplification techniques may then be used on the subset to produce a reduced-sample current subset of data values that are converted to an angular coordinate system. A space partitioning data structure such as a k-dimensional tree that stores converted reduced-sample previous subsets of the data stream may then be traversed to determine one or more nearest neighbors to the current subset. The predicted next sample for the data stream may be generated from the nearest neighbors. The space partitioning data structure may be updated to include the current subset, and the process may be repeated with a new current subset.


