The present disclosure relates to methods, systems, and apparatuses for featurizing time-
series data to enhance
machine learning model training. Time-
series data, such as transaction records, is preprocessed to identify fields including descriptive and categorical information. Categories are assigned using a first
machine learning model, and tags are applied based on domain-specific patterns or large language models. The processed data is organized into a
star schema data structure comprising a fact structure and associated dimension structures. Features are generated from the
data structure based on
time windows, incorporating statistical
metrics and identified patterns. These features are provided to a
machine learning module to
train a second
machine learning model, improving accuracy and adaptability for applications such as customer behavior prediction and financial analysis. The disclosed approach addresses challenges of
high dimensionality,
noise, and temporal dependencies in time-
series data, enabling robust and contextually
relevant feature generation.