Time-Series Record Classification Using Random Re-Censoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Artificial intelligence models for classifying time-series data are hindered by inherent bias due to the reliance on outdated data, which is not reflective of current circumstances, leading to inaccurate predictions and model inefficiencies.
Innovation Solution
Implementing random data re-censoring to purposefully make training data incomplete by censoring events based on proximity to a set build date, thereby mitigating bias and ensuring more recent data has a larger impact on model curves.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If complete historical data is used for training, then model accuracy is improved, but bias from outdated data increases
Solution Approach 1:
The patent extracts and removes biased outdated data from the training set by applying random censoring to historical records. This selectively extracts problematic data points while retaining useful information, resolving the contradiction between using complete data for accuracy and removing biased data for reliability.
Solution Approach 2:
The patent changes the parameter of data completeness by intentionally introducing randomness and incompleteness through censoring mechanisms. This parameter change allows the model to learn from incomplete patterns while reducing bias, simultaneously improving both accuracy and reliability.
2Stability of the object's composition
If older data is weighted more heavily, then model stability is improved, but adaptability to current circumstances deteriorates
Solution Approach 1:
The patent introduces dynamic weighting through random censoring that varies by data age and outcome realization status. This dynamic approach allows the model to adaptively balance between stable historical patterns and current circumstances, improving both stability and adaptability simultaneously.
Solution Approach 2:
The patent applies periodic re-censoring and retraining cycles that systematically update the model with fresh data while periodically removing outdated biased records. This periodic action maintains model stability through consistent training while improving adaptability to changing conditions.
3Reliability
If more complete training data is used, then model performance is improved, but processing time increases
Solution Approach 1:
The patent extracts only the essential training signals needed for model performance by censoring redundant and biased historical data. This extraction reduces the volume of processing required while maintaining model performance, resolving the time-performance tradeoff.
Solution Approach 2:
The patent discards outdated biased data through random censoring while recovering and emphasizing more relevant recent patterns. This selective discarding and recovering reduces processing time by eliminating unnecessary data while maintaining model performance.
Data Source
AI summary
Methods and systems are described herein for improving data processing efficiency of classifying user files in a database. More particularly, methods and systems are described herein for improving data processing efficiency of classifying user files in a database in which the user files have a temporal element. The methods and systems described herein accomplish these improvements by mitigating bias using random data censoring.


