Synthetic Time-Series Data Generation Through Multi-Scale Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional systems and methods for generating synthetic time-series data are limited in generating realistic, multi-dimensional data across various time scales and dimensions, often requiring human judgment for data distribution selection and being restricted to one-directional or limited time frames.
Innovation Solution
A system and method utilizing machine learning to optimize segment parameters and distribution measures, involving the training of parameter and distribution models to generate synthetic datasets from time-series data, enabling flexible and accurate multi-dimensional data generation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional systems generate synthetic time-series data, then data can be produced for confidentiality or availability purposes, but the data lacks realism and cannot capture changes over time accurately
Solution Approach 1:
The patent segments time-series data into multiple data segments at different time scales (e.g., hourly, daily, weekly segments). Each segment is processed independently to learn temporal patterns at that specific scale, then combined to generate realistic synthetic data that captures changes across multiple time horizons, resolving the limitation of conventional single-scale approaches
Solution Approach 2:
The patent adds the time scale dimension by training separate models for different temporal resolutions (short-term, medium-term, long-term patterns). This multi-dimensional approach allows the system to capture complex temporal dynamics that single-scale models miss, improving realism without requiring exponentially more computational resources
2Adaptability or versatility
If conventional approaches use pre-defined data distributions, then data generation is simpler, but human judgment is required and flexibility is reduced
Solution Approach 1:
The system automatically selects and adapts data distributions for each segment without human intervention. The model learns the appropriate distribution characteristics from the training data and applies them autonomously during synthesis, eliminating the need for manual distribution selection while maintaining flexibility to handle diverse data types
3Reliability
If conventional systems generate synthetic data within observed parameter ranges, then data stays within known bounds, but the data lacks extrapolation capability and realism
Solution Approach 1:
The patent employs dynamic parameter generation where segment parameters (min, max, mean, variance) are learned from training data and applied adaptively to each generated segment. This allows the synthetic data to naturally extend beyond observed ranges while maintaining statistical consistency, capturing realistic variations and anomalies that static parameter approaches cannot produce
4Adaptability or versatility
If conventional methods generate time-series data in one direction only, then the generation process is simpler, but the data cannot represent multi-directional temporal patterns
Solution Approach 1:
The patent creates a universal generation framework that can produce synthetic data in multiple temporal directions (forward, backward, bidirectional) using the same trained model. The segment-based architecture allows the system to flexibly generate data sequences in any direction by simply changing the generation order, making the model multi-functional without requiring separate models for each direction
Data Source
AI summary
Systems and methods for generating synthetic data are disclosed. For example, a system may include one or more memory units storing instructions and one or more processors configured to execute the instructions to perform operations. The operations may include receiving a dataset including time-series data. The operations may include generating a plurality of data segments based on the dataset, determining respective segment parameters of the data segments, and determining respective distribution measures of the data segments. The operations may include training a parameter model to generate synthetic segment parameters. Training the parameter model may be based on the segment parameters. The operations may include training a distribution model to generate synthetic data segments. Training the distribution model may be based on the distribution measures and the segment parameters. The operations may include generating a synthetic dataset using the parameter model and the distribution model and storing the synthetic dataset.


