Data Unfolding via Temporal Segmentation for Prediction Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Insufficient data sets limit the accuracy of data analytics, particularly in predicting events like hospitalization, due to the lack of robust data points, leading to inaccurate results when applied to larger populations.
Innovation Solution
The method of data folding and unfolding generates multiple records from existing data points by dividing time intervals, allowing for the creation of a more extensive and robust data set that can be used to train prediction models, increasing data sampling and accuracy through machine learning techniques.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data folding and unfolding is applied to generate multiple records from existing data points, then the data set size increases, but the risk of introducing bias increases
Solution Approach 1:
The patent applies segmentation by dividing the time interval into multiple sub-intervals (e.g., splitting a 30-day period into weekly segments). Each segment becomes an independent data point with its own event indicator, allowing the model to learn temporal patterns while maintaining data integrity. This resolves the contradiction by increasing data quantity through temporal segmentation rather than artificial duplication.
Solution Approach 2:
The patent introduces a temporal dimension by creating multiple records from single events based on different time intervals. Instead of duplicating data points, it unfolds time-based information into multiple dimensional representations (e.g., event occurred in week 1, week 2, week 3, week 4), increasing data set size while preserving the original information structure and avoiding bias.
2Measurement precision
If more data points are collected to improve prediction accuracy, then the accuracy of data analytics increases, but the complexity of data collection and processing increases
Solution Approach 1:
The patent applies preliminary action by pre-processing and structuring data into standardized formats with consistent time intervals and event indicators before model training. This preliminary organization allows the system to generate multiple analytical records from existing data without requiring additional complex data collection infrastructure, thereby improving prediction accuracy while maintaining manageable complexity.
Solution Approach 2:
The patent uses copying by creating multiple analytical representations of existing data points through time-based segmentation. Rather than collecting new data from additional sources, it generates multiple copies of the same underlying data with different temporal perspectives, increasing data utility without expanding data collection complexity.
3Quantity of substance
If time intervals are divided to generate more records, then the data sampling increases, but the processing time increases
Solution Approach 1:
The patent segments time intervals into manageable units (e.g., daily, weekly, monthly) to generate multiple records from single events. This segmentation strategy increases data sampling by creating granular time-based observations while keeping processing time manageable through efficient computational algorithms that can handle segmented temporal data.
Solution Approach 2:
The patent applies parameter changes by allowing flexible adjustment of time interval granularity (e.g., changing from daily to weekly segments). This enables optimization of the trade-off between data sampling quantity and processing time based on specific analytical needs, resource availability, and event frequency characteristics.
Data Source
AI summary
Systems and methods for data unfolding are disclosed. For example, it may be desirable or necessary to increase a data set, such as for increasing accuracy of one or more predictive models. Data set proliferation without introducing unnecessary bias may be important for increasing such accuracy. Described herein are system and methods that allow for data set proliferation by generating records based on whether an event occurred with respect to an entity during multiple time intervals. A record may be generated for each time interval and the associated data may be unfolded and disassociated, at least partly, from other records related to the entity. Those records may then be used for data analytics and/or predictive model generation, for example.


