Time Series Event Prediction via Data Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for predicting events based on time series data often rely solely on the original data set, lacking the ability to accurately compare and combine data from multiple sources to enhance prediction accuracy.
Innovation Solution
A system that generates time series data sets from various sources, extracts subsets of contiguous data points, groups them into clusters, and uses a machine learning model to determine the probability of an event occurrence, while also comparing the data sets for similarity and transforming scores into a new distribution.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing methods rely solely on the original data set for prediction, then the prediction process is simple, but the prediction accuracy is insufficient
Solution Approach 1:
The patent segments the original time series data into multiple subsets of contiguous data points. This segmentation allows the system to analyze different portions of the data independently and combine their predictions, thereby improving overall prediction accuracy while maintaining a structured approach to handling complexity
Solution Approach 2:
The patent merges predictions from multiple data subsets by grouping them into clusters and combining cluster predictions. This merging process integrates information from different data segments, enhancing prediction accuracy through comprehensive data utilization while systematically managing the complexity through structured combination methods
2Reliability
If data from multiple sources are integrated, then prediction reliability improves, but data processing complexity increases
Solution Approach 1:
The patent performs preliminary actions by pre-processing and organizing data from multiple sources into structured time series data sets before main analysis. This includes extracting contiguous data point subsets and preparing them for clustering, which simplifies subsequent processing and enhances reliability through thorough preliminary data preparation
Solution Approach 2:
The patent introduces clustering as an intermediary step between raw multi-source data and final predictions. The clustering process groups similar data patterns together, acting as a mediator that simplifies the integration of diverse data sources while maintaining prediction reliability through systematic pattern recognition and grouping
Data Source
AI summary
Some embodiments provide a non-transitory machine-readable medium that stores a program. The program receives data from a set of data sources. Based on the data, the program further generates a time series data set that includes a plurality of data points. The program also determines a set of subsets of data points from the time series data set. The program further groups the set of subsets of data points into a set of clusters. The program also provides the set of clusters as input to a machine learning model to cause the machine learning model to determine a probability of an occurrence of an event based on the set of clusters.


