Time Series Prediction Using Space Partitioning Data Structures

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current prediction techniques for data streams become computationally expensive and slow with large volumes of data, often resulting in delayed or ineffective predictions, especially in systems with high activity levels, such as database systems experiencing peak loads.

Innovation Solution

The proposed solution involves a prediction system that uses polyline simplification to reduce sample subsets, converts data to angular coordinates, and stores them in a space partitioning data structure like a k-dimensional tree, allowing for faster prediction times by identifying nearest neighbors and generating predicted samples efficiently.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If current prediction techniques are used on large volumes of data, then prediction accuracy may be maintained, but processing time becomes excessively long and predictions become delayed

Engineering Contradiction:
Improveprediction accuracyVSAvoidprediction delay
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The data stream is divided into fixed-size sliding windows that partition historical data into manageable subsets. Each window contains a fixed number of recent data points, allowing the system to process only relevant recent history rather than the entire data stream. This segmentation enables efficient computation while maintaining prediction accuracy by focusing on the most relevant recent patterns.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts only the essential features from historical data by using fixed-size windows that capture recent patterns without storing entire historical sequences. By extracting and processing only the necessary recent data points within each window, the system avoids the computational burden of processing all historical data while maintaining the ability to make accurate predictions.

Inventive Principle:
Principle #2Taking out (Extraction)

2Reliability

If regression techniques are applied to analyze large data streams, then prediction capability is maintained, but computational complexity increases significantly

Engineering Contradiction:
Improveprediction capabilityVSAvoidcomputational complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The data is segmented into fixed-size sliding windows, which limits the amount of historical data processed for each prediction. This segmentation reduces computational complexity by ensuring that regression analysis is performed only on a manageable subset of recent data points rather than the entire historical data stream, while still maintaining prediction capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary data organization by maintaining pre-structured sliding windows of historical data in memory. This preliminary organization allows the regression algorithm to quickly access and process only the necessary recent data points without having to retrieve and organize large volumes of historical data during prediction, thereby reducing computational complexity.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11315036B2Prediction for time series data using a space partitioning data structure
Publication Date: 2022.04.26 PAYPAL INC
  • US11315036B2 patent drawing
  • US11315036B2 patent drawing
  • US11315036B2 patent drawing

AI summary

Techniques are disclosed for a computer system to predict a next sample for a data stream that specifies data values of one or more variables. A current subset of data values and previous subsets of data values is determined, and polyline simplification techniques may then be used on the subset to produce a reduced-sample current subset of data values that are converted to an angular coordinate system. A space partitioning data structure such as a k-dimensional tree that stores converted reduced-sample previous subsets of the data stream may then be traversed to determine one or more nearest neighbors to the current subset. The predicted next sample for the data stream may be generated from the nearest neighbors. The space partitioning data structure may be updated to include the current subset, and the process may be repeated with a new current subset.