System for time series generation (TSG) model selection
The system addresses the challenges of TSG model selection by using LLMs and RAG for interactive model evaluation, providing tailored recommendations and comprehensive metrics, enhancing the practical application of TSG methods in diverse industries.
Patent Information
- Application Number
- PCT/SG2025/050490
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-22
- Filing Date
- 2025-07-18
- Publication Date
- 2026-01-29
AI Technical Summary
Existing time series generation (TSG) methods face challenges in practical application due to the variation between methods, lack of standardized evaluation, and the need for extensive technical knowledge to select suitable models for specific industrial applications, hindered by a lack of dynamic platforms for comparing performance in real-world data contexts.
A system integrating Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) to facilitate user interaction, providing a comprehensive and systematic evaluation of TSG models through an interactive platform that assists in selecting suitable models based on specific needs and industrial applications, using evaluation metrics tailored to data characteristics and requirements.
Enables effective and efficient selection of TSG models by bridging user intent with appropriate model configurations and domain-specific guidance, offering personalized recommendations and thorough evaluation, enhancing the practical application of TSG methods in various industries.
Smart Images

Figure SG2025050490_29012026_PF_FP_ABST
Abstract
Description
SYSTEM FOR TIME SERIES GENERATION (TSG) MODEL SELECTIONCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims the benefit of priority to the Singapore application no. 10202402169R filed 22 July, 2024, the contents of which are hereby incorporated by reference in their entirety for all purposes.TECHNICAL FIELD
[0002] This application relates generally to the field of data generation, and more particularly, to a system for time series generation (TSG) model selection.BACKGROUND
[0003] Time series generation (TSG) methods find application in various industries in providing generated synthetic data that mirrors or mimics real-world characteristics. The essence of TSG lies in the ability to generate synthetic data that retains the intrinsic statistical properties and temporal dependencies of the original or real-world data. This provided application in domains such as finance for risk assessment, environmental science for climate modelling, and manufacturing for predictive maintenance, where real data is limited, sensitive, and / or unevenly distributed.
[0004] While TSG methods have gained acceptance in various industries, the practical applications of TSG methods have not been straightforward. This is due to the variation between TSG methods and the associated strengths and limitations. Moreover, assessing the quality of generated time series necessitates a suite of thorough measures that scrutinize various facets of data fidelity and temporal coherence.SUMMARY
[0005] According to an aspect, disclosed herein a system. The system comprises: memory storing instructions; and a processor coupled to the memory and configured to process the stored instructions to implement: a module configured to perform a method of Time Series Generation (TSG) model selection. The method including: receiving a user prompt, the user prompt comprising a user input and a user-provided TSG dataset; providing the user input as input to a first machine learning model, to determine the user prompt as a TSG query; providing the user input and the user-provided TSG dataset as input to a second machine learning model, to select at least one shortlisted TSG model from a TSG database; and using the first machine learning model, providing the at least one shortlisted TSG model as a response to the TSG query.
[0006] According to another aspect, disclosed herein a portable device comprising the system as described above.BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Various embodiments of the present disclosure are described below with reference to the following drawings:FIG. 1 is a schematic diagram showing a system and a device for time series generation model selection according to embodiments of the present disclosure;FIG. 2 is a schematic diagram showing a system and a device for time series generation model selection according to various embodiments,FIG. 3 is a schematic diagram illustrating a system for time series generation model selection according to embodiments of the present disclosure,FIG. 4 is a schematic diagram illustrating another system for time series generation model selection according to embodiments of the present disclosure;FIG. 5 is a schematic diagram illustrating another system for time series generation model selection according to embodiments of the present disclosure,FIG. 6 is a schematic diagram illustrating another system for time series generation model selection according to embodiments of the present disclosure;FIG. 7 is a schematic diagram illustrating another system for time series generation model selection according to embodiments of the present disclosure;FIG. 8 is a schematic diagram illustrating another system for time series generation model selection according to embodiments of the present disclosure;FIG. 9 is a flow chart illustrating a method of time series generation model selection according to embodiments of the present disclosure;FIG. 10 is a schematic diagram illustrating a system for time series generation model selection according to an exemplary embodiment;FIG. 11 is a schematic diagram illustrating a data flow for the system of FIG. 10;FIGs. 12A and 12B are screenshots of an implementation of the system of FIG. 10;FIGs. 13 A and 13B are experimental results from the implementation of FIGs. 12A and 12B, FIG. 14 is a schematic diagram of a processor system.DETAILED DESCRIPTION
[0008] The following detailed description is made with reference to the accompanying drawings, showing details and embodiments of the present disclosure for the purposes of illustration. Features that are described in the context of an embodiment may correspondingly be applicable to the same or similar features in the other embodiments, even if not explicitly described in these other embodiments. Additions and / or combinations and / or alternatives as described for a feature in the context of an embodiment may correspondingly be applicable to the same or similar feature in the other embodiments.
[0009] In the context of various embodiments, the articles “a”, “an” and “the” as used with regard to a feature or element include a reference to one or more of the features or elements.
[0010] In the context of various embodiments, the term “about” or “approximately” as applied to a numeric value encompasses the exact value and a reasonable variance as generally understood in the relevant technical field, e.g., within 10% of the specified value.
[0011] As used herein, the term “and / or” includes any and all combinations of one or more of the associated listed items.
[0012] As used herein, terms “concurrently”, “simultaneously”, “at the same time”, or the like, may refer to events or actions that coincide or overlap within a period of time, regardless of whether the events start at the same time instant, and regardless of whether the events end at the same time instant.
[0013] As used herein, the term “Time Series Generation (TSG) query” may be used interchangeably with the terms “inquiry”, “request”, “question”, etc., and may generally refer to an enquiry from a user seeking information relating to TSGmodel(s), as it may be understood. In an exemplary TSG query, a user may be seeking a selection of TSG models suitable for an intended technical domain and / or application. In another exemplary TSG query, a user may be seeking general overviews of different TSG models suitable for generating dataset similar to a specific historical dataset. In yet another exemplary TSG query, a user may be seeking a quantitative comparison between two or more TSG models for an intended technical domain.
[0014] As used herein, the term “user-provided TSG dataset” may generally refer to a time series dataset received by the proposed system and / or device, typically provided as input from a user. However, it may be noted that the user-provided TSG dataset may correspond to real- life or historical dataset obtained from actual data, but is not limited there to In some instances, the user-provided TSG dataset may also be generated TSG dataset previously generated from one or more TSG models.
[0015] A motivation for Time Series Generation (TSG) is to generate synthetic time series data that is similar or even indistinguishable from the corresponding real-world data, while concurrently preserving the underlying patterns and correlations. TSG plays an important role in applications such as anomaly detection and privacy preservation. Some exemplary approaches in TSG may include different types of TSG models or methods, such as: Generative Adversarial Networks (GANs), Variational AutoEncoders (VAEs), and Flow-based models. For example, GAN-based models, like TimeGAN and RTSGAN, may capture temporal dependencies through adversarial training. VAE-based methods, such as TimeVAE and TimeVQVAE, may generate new time series using latent representations, ensuring data fidelity. Flow-based models may utilize transformations to effectively model complex data distributions.
[0016] However, translation of TSG methods into industry applications or practices face several challenges. There remains a cognitive gap for users where the sheer volume of TSG methods and metrics may be daunting. In addition, the understanding of the various nuanced differences and suitability in TSG methods often necessitate extensive technical knowledge on each TSG model rendering a comprehensive comparison tedious. Hence, identifying the suitable TSG model(s) for specific industrial application(s) remain challenging. Further, the lack of a dynamic platform hinders industry users from effectively exploring or comparing the performance of various TSG models in real-world data contexts.
[0017] The proposed system and method address the above limitations regarding TSG method selection and application The proposed system may be configured as an interactive system for Time Series Generation (TSG) model or method selection. The proposed system and method may provide a consistent and systematic evaluation for TSG models / methods based on an interaction or a series of interactions with a user. The proposed system and method mayfurther assist in the selection or determination of suitable TSG models / methods based on specific needs, requirements or specific industrial applications.
[0018] In addition, the proposed system and method may comprise evaluation measures or evaluation matrices tailored for specific data characteristics and requirements for the technical domain or application. Moreover, assessing the quality of generated time series necessitates a suite of thorough measures that scrutinize various facets of data fidelity and temporal coherence. This represents a significant leap in standardized TSG evaluation, offering a benchmarking framework with a broad range of comprehensive metrics. Apart from selecting or shortlisting appropriate TSG models / methods, the proposed system / method may also provide evaluation measures based on each of the shortlisted TSG models / methods. Further, the proposed system / method may perform evaluation on each of the shortlisted TSG models / methods based on evaluation measures tailored to specific data characteristics and requirements. In addition, visualization tools may also be provided in enabling a comprehensive and clear comparison between shortlisted TSG models.
[0019] The proposed system and method integrate one or more Large Language Models (LLMs) which are often useful in understanding and generating natural language to facilitate user interaction. Trained on vast text corpora, LLMs excel in grasping intricate linguistic structures and semantic nuances The LLMs interact with users through natural language prompts, accommodating a broad range of applications. Importantly, LLMs may engage in multi-round dialogues, taking into account prior interactions for more contextualized and precise responses. This feature is especially useful for the proposed system, evolving the proposed system from a rule-based system to an automated conversational agent. In practice, this means users can interactively seek TSG method recommendations for specific datasets, moving beyond manual searches in tables. LLMs dynamically tailor responses to the user’s specific requirements and context.
[0020] Further, the capabilities of LLMs may be further augmented by Retrieval- Augmented Generation (RAG). This enables the proposed system to harness external knowledge bases during the generation process, thus allowing pertinent information to be obtained in response to user queries. In the RAG process, upon receiving a user prompt (or a user query), the system may search a related knowledge base, extract relevant data, and integrate the selected data into the response to the query to be provided via the LLM. The integration of RAG with LLM enhances the response with detailed, accurate information and reduces errors or hallucinations in the content.
[0021] In an exemplary operation of the proposed system and method, if a user asks about a specific TSG method, the RAG-equipped or RAG-integrated LLM(s) may retrieve detailed information from a knowledge base, which enhances and validates user interactions. The depth of TSG related knowledge and support in decision-making allows RAG- integrated LLM(s) to be suitably deployed for TSG applications and / or selection.
[0022] Referring to FIGs. 1 and 2, in one aspect, the proposed system 100 for TSG model selection according to various embodiments of the present disclosure. The system 100 may comprise a processor 901 configured for perfor ing a method of Time Series Generation (TSG) model selection. In avoidance of doubt, in some implementations, the system 100 and the processor 901 may be located in a computing server remote from a user 80. In such scenarios, the user 80 may access the system 100 / processor 901 via a network connection using a local computing device.
[0023] In other embodiments, the processor 901 may be disposed locally on a device 1 10, such as a desktop computer. In various embodiments, the device 110 may also include a user interface 902 and an input device 935 operable by a user 80. Referring to FIG. 1, the device 110 may include a screen for displaying the user interface 902, and a keyboard as the input device 935. In various embodiments, referring to FIG. 2, the device 110 may be configured as aportable device, such as a smart device. The portable device 110 may include a touch screen for both displaying the user interface 902 and acting as the input device 935.
[0024] FIG. 3 is a schematic diagram illustrating a structure of the system 100 according to various embodiments of the disclosure. Additionally, FIG. 3 also illustrates dataflow for a method of TSG model selection performed by the system 100 with the various signal connections and signal flows in the system 100. The system 100 may include a user interface 902; a first machine learning model 210; a second machine learning model 220, and a Time Series Generation (TSG) database 230. The user interface 902 may be in signal communication with the first machine learning model 210 and the second machine learning model 220. The first machine learning model 210 may be in signal communication with the second machine learning 220. The second machine learning 220 may be in signal communication with the TSG database 230. The TSG database 230 may act as a knowledge base to the system 100. The TSG database 230 may be configured to store various TSG related data, such as a plurality of TSG models, a plurality of pre-stored TSG datasets, a plurality of evaluation measures, a plurality of visualization tools, etc. The TSG database 230 may be a private database dedicated to TSG related data, instead of a publicly available database.
[0025] The heterogeneity and complexity of time series generation tasks demand not only powerful models but also effective ways to bridge user intent with appropriate model configurations, evaluation metrics, and domain-specific guidance. In various embodiments, the first machine learning model 210 may be a large language model (LLM) and the second machine learning model 220 may be a Retrieval-Augmented Generation (RAG) model The RAG may be configured to retrieve or obtain information from the TSG database 230. Collectively, the LLM 210, the RAG 220 and the TSG database 230 may serve as an intelligent assistant to guide users through the modelling process, backed by a structured knowledge base.
[0026] Applicability of TSG models often depends on nuanced dataset characteristics (e.g., length, dimensionality, periodicity), use-case requirements (e.g., forecasting, classification), and technical constraints (e.g., computational efficiency). However, this mapping between model choice and task context remains tacit in most systems. Moreover, end users, especially those from applied domains, frequently lack the expertise or time to navigate through research papers or empirical benchmarks. The RAG model tills this gap by offering natural language support through retrieval -enhanced generation.
[0027] According to various embodiments, the system 100 (comprising the first machine learning model 210 and the second machine learning model 220) may be trained collectively using a variety of training datasets obtained from different technical domains. Alternatively, the first machine learning model 210 and the second machine learning model 220 may first be independently trained prior to being collectively trained. Preprocessing may be performed on the training dataset to transform long time series into fixed length overlapping windows. The training dataset may be split into training dataset and validation dataset. The training dataset may also be stored in the TSG database 230 (or knowledge database) for repeated training or evaluation runs.
[0028] According to various embodiments, to perform a Time Series Generation (TSG) model selection, the system 100 may be configured to receive a user prompt 310 via the user interface 902 and the input device 935. The user prompt 310 may correspond to an active query from a user 80 to the system 100. The user prompt 310 may comprise a user input 312 and a user-provided Time Series Generation (TSG) dataset 314. The user input 312 may comprise one or a series of natural language input from a user 80 using the input device 935. In various embodiments, the user input 312 may correspond to a series of interactions or natural language exchanges between the system 100 and the user 80. The user input 312 may also include contextual information from the user 80, such as specific requirements for the TSG model(s)and / or a relevant technical domain for the TSG model(s). The user-provided TSG dataset 314 may be a time series dataset which is relevant to the user prompt 310. As such, mere interaction between the user 80 and the system 100 without the provision of the user-provided TSG dataset 314 does not constitute the user prompt 10.
[0029] In various embodiments, the system 100 may support raw time series dataset (e.g. standard formats such as .csv and pkl), as well as benchmark datasets from the UCR / UEA archive. The raw time series datasets may comprise continuous or discretized real-valued data points across multiple sensors or channels. If there are any missing values, the missing values may be linearly interpolated along the temporal axis to ensure data continuity.
[0030] In various embodiments, the user-provided TSG dataset 314 may be used to determine a nature or context of the user prompt 310. For example, the user-provided TSG dataset 314 may be indicative of a technical domain, such as in weather prediction, in which the user 80 finds relevance and wishes to obtain related time series data.
[0031] Response to receiving the user prompt 310, the user input 312 may be provided to the first machine learning model 210 as input. The first machine learning model 210 may be a generative machine learning model, such as a large language model (LLM), trained to interact with the user 80. Based on the user input 312, the first machine learning model 210 may determine the user prompt 310 as a Time Series Generation (TSG) query 320. The TSG query 320 may correspond to an enquiry from the user 80 seeking information relating to TSG models and / or TSG methods.
[0032] In various embodiments, the system 100 may be provided with a prompt construction module for extracting metadata from the user input 312 and / or the user-provided TSG dataset 314. The prompt construction module may be configured as an independent module or be integrated into one or both of the first and second machine learning models 210 / 220. Based on the extracted metadata, a structured natural language prompt may be assembled to guide theusers. The prompt construction module may provide a concise and human-readable summary of the user-provided TSG dataset 314, while hiding low-level technical terms that may overwhelm non-experts.
[0033] Upon determining the user prompt 310 as a TSG query 320, or in other words, determining a request pertaining to TSG from the user 80, the user input 312 may be provided as input to a second machine learning model 220. Concurrently, the user-provided TSG dataset 314 may also be provided as input to a second machine learning model 220. Based on the user input 312 and the user-provided TSG dataset 314, the second machine learning model 220 may select one or more shortlisted TSG model(s) 330 from the TSG database 230. Example of the shortlisted TSG model(s) 330 may include various categories of TSG models such as: Generative Adversarial Network (GAN)-based Models, Variational AutoEncoders (VAE)- based Models, Flow-based Models, and Mixed-Type Models. The above corresponds to the system 100 identifying the user prompt 310 as a TSG query 320, and hence proceeds to a subsequent step of selecting the shortlisted TSG model(s) 330 which corresponds to the user prompt 310.
[0034] In various embodiments, the user input 312 may comprise an indication of a nature of the TSG query 320, such as a context 340 of the user prompt 310. For example, the user input 312 may include a relevant technical domain in which the user 80 expressed interest in. In another example, the user input 312 may include a specific nature of the request, such as providing one or more evaluation measures on shortlisted TSG model(s) 330. Similarly, the system 100 may also determine the context 340 from the user-provided TSG dataset 314. As such, the context 340 may comprise information obtained from one or both of the user input 312 and the user-provided TSG dataset 314.
[0035] In addition, the context 340 may also be provided to the second machine learning model 220 to aid in the selection of the shortlisted TSG model(s) 330 from the TSG database230. Hence, the second machine learning model 220 may retrieve the shortlisted TSG model(s)330 from the TSG database 230 based on the user input 312, the user-provided TSG dataset 314 and the context 340.
[0036] In various embodiments, upon selecting the shortlisted TSG model(s) 330, the system 100 may provide the shortlisted TSG model(s) 330 as a response 350 to the TSG query 320. The system 100 may provide the shortlisted TSG model(s) 330 as a response 350 to the TSG query 320 using the first machine learning model 210. In other words, the system 100 may provide the shortlisted TSG model(s) 330 via one or more natural language interactions with the user 80. In various embodiments, in response to the user prompt 310 the system 100 may provide or display the one or more shortlisted TSG model(s) 330 via the user interface 902.
[0037] As the proposed system 100 may be interactive in nature, the process of providing the shortlisted TSG model(s) 330 may be an iterative process in which the system 100 engages and interacts with the user 80 in addressing the needs and requirements of the user 80.
[0038] In various embodiments, upon the system 100 providing the one or more shortlisted TSG model(s) 330 and responsive to a further user input, such as a comment on the shortlisted TSG model(s) 330, the system 100 may provide one or more subsequent shortlisted TSG model(s) based on the further user input and the user-provided TSG dataset 314. This enables the system 100 to refine the shortlisted TSG model(s) 330 based on feedbacks from the user 80. Hence, upon refining, the one or more subsequent shortlisted TSG model(s) may be nonidentical to the shortlisted TSG model(s) 330 previously provided by the system 100.
[0039] In various embodiments, the system 100 may provide the further user input to the first machine learning model 210 and the user-provided TSG dataset 314 to the second machine learning model 220 to refine the context 340 of the TSG query 320. With the refined context, the system 100 may provide one or more subsequent shortlisted TSG model(s). Further, inaddition to the further user input, a subsequent user-provided TSG dataset may also be provided to refine the context 340 of the TSG query 320.
[0040] The TSG database 230 may comprise or store a plurality of TSG models. The plurality of TSG models may include multiple types of TSG models, such as: Generative Adversarial Network (GAN)-based Models, Variational AutoEncoders (VAE)-based Models, Flow-based Models, and Mixed-Type Models. Exemplary GAN-based Models may include RGAN, TimeGAN, RTSGAN, COSCI-GAN, AEC-GAN, etc. Exemplary VAE-based Models may include TimeVAE, TimeVQVAE, etc. Flow-based Models often utilize explicit likelihood models or Ordinary Differential Equations (ODEs), and typically feature coupling layers, allow for a computable Jacobian determinant and reversibility. This is further enhanced by specific transformation techniques crucial for modelling complex data distributions. Exemplary Mixed- Type Models may include Fourier Flows, GT-GAN, LS4, etc.
[0041] Further referring to FIG. 4, the system 100 may further comprise a data processing module 215. The data processing module 215 may be configured to preprocess the user- provided TSG dataset 314 to obtain a preprocessed TSG dataset 315. The preprocessed TSG dataset 315 may be used to determine a context 340 of the TSG query 320. The context 340 of the TSG query 320 may comprise a domain-specific application of the user-provided TSG dataset 314.
[0042] In various embodiments, the data processing module 215 may normalize and / or segment the user-provided TSG dataset 314 to obtain the preprocessed TSG dataset 315. The preprocessing method may include transforming long time series into fixed-length overlapping windows. If the desired sequence length is not explicitly provided, the sequence length may be inferred automatically using autocorrelation analysis. In various embodiments, the data processing module 215 may obtain preprocessing methods from the TSG database 230.
[0043] A robust and consistent data preprocessing pipeline enables the effectiveness of time series generation models. A standardized data input and preprocessing procedure is disclosed, which transforms raw multivariate time series data into normalized, fixed-length segments suitable for training and evaluation. In an example, the system 100 may support two types of time series datasets: (1) raw time series files in standard formats such as csv and pkl, and (2) benchmark datasets from the UCR / UEA archive. The raw datasets may comprise continuous or discretized real-valued observations across multiple sensors or channels. Missing values, if any, may be linearly interpolated along the temporal axis to ensure data continuity.
[0044] In exemplary embodiments, user-provided TSG dataset 314 with long time series may be transformed into fixed-length overlapping windows. If the desired sequence length is not explicitly provided, the sequence length is inferred automatically using autocorrelation analysis. Specifically, for each channel, an autocorrelation function (ACF) up to min(20000, L) samples are computed using the Fast Fourier Transform (FFT). The first local maximum of the ACF beyond lag 3 is identified to estimate the dominant periodicity. The inferred sequence length is set to the mean of all valid periods across channels, or 125 if no valid local maxima are found. This procedure adapts dynamically to varying temporal patterns across different datasets.
[0045] After determining the sequence length, a non-overlapping sliding window of size T is applied to segment the input into shorter sequences. Formally, for an input matrix X 6 IR / 'xCwith L timesteps and C channels, to construct a 3D tensor X £ j^vxrxc,wherejy = £ — T + 1 is the number of generated windows.
[0046] Similar preprocessing process may be applied to training datasets during model training. To reduce temporal correlation between training dataset examples and promote generalization, all windowed sequences may be randomly shuffled. The data is then split into training and validation sets using a default validation ratio of 10%. If enabled, a Min-Maxnormalization is applied independently for each feature based on the dataset distribution. The same transformation is then applied to the validation data to avoid data leakage. This enables the input features to lie within a bounded range (typically [0,1]), which is important for the stability of many deep generative models. To support efficient reusability, all preprocessed data are serialized and stored in compressed pkl format. The saved data can be loaded later via a dedicated utility, bypassing the preprocessing stage during repeated training or evaluation runs.
[0047] Further referring to FIGs. 5 and 6, the system 100 may further comprise a benchmarking module 400 in signal communication between the first machine learning model 210 and the user interface 902. The benchmarking module 400 may comprise a simulation module 410 in signal communication between the first machine learning model 210 and the user interface 902. The simulation module 410 may be configured to perform simulation based on each of the one or more shortlisted TSG model(s) 330. The simulation module 410 may be configured to simulate and provide one or more generated TSG data 345 based on the one or more shortlisted TSG model(s) 330. Hence, it may be appreciated that each of the generated TSG data 345 may correspond to a respective one of the one or more shortlisted TSG model(s) 330. The generated TSG data 345 may act as a response 350 to the TSG query 320.
[0048] In various embodiments, referring to FIG. 5, responsive to the user prompt 310 and using the simulation module 410, the system 100 may input or provide a pre-stored TSG dataset 365 from the TSG database 230 to each of the one or more shortlisted TSG model(s) 330 to generate a respective generated TSG data 345. The TSG database 230 may comprise a plurality of pre-stored TSG datasets 365 In other words, the one or more shortlisted TSG model(s) 330 are tested / used for generating TSG data based on the pre-stored TSG datasets 365.
[0049] Exemplary pre-stored TSG datasets 365 may include but is not limited to: i) Dodgers Loop Game (DLG) comprising loop sensor data from the Glendale on-ramp for the 101 North freeway in Los Angeles; ii) Historical stock data from 2004 to 2019, including volume andhigh, low, opening, closing, and adjusted closing prices; iii) Stock Long which corresponds to the stock dataset but with a sequence length of 125; iv) Exchange which comprises daily currency exchange rates of eight countries (i.e., Australia, Britain, Canada, Switzerland, China,Japan, New Zealand, and Singapore) from 1990 to 2016; v) Energy use in a low-energy building; Energy Long which corresponds to energy use dataset but with a sequence length of 125; vi) Electroencephalography (EEG) data which aids in understanding brainwave patterns, especially those under specific cognitive conditions or stimuli; (vii) HAPT which comprises recordings of 30 subjects performing activities of daily living captured via waist-mounted smartphones with embedded inertial sensors; (viii) Air quality, meteorological, and weather forecast data from 4 major Chinese cities: Beijing, Tianjin, Guangzhou, and Shenzhen from 2014 / 05 / 01 to 2015 / 04 / 30; (ix) sensor data from three boilers from 2014 / 03 / 24 to 2016 / 1 1 / 30 to monitor the operating states. The pre-stored TSG datasets 365 may correspond to various technical domains, such as: traffic, financial, appliances, medical, sensor, industrial, but is not limited thereto.
[0050] In alternative embodiments, referring to FIG. 6, responsive to the user prompt 310 and using the simulation module 410, the system 100 may input or provide the user-provided TSG dataset 314 to each of the one or more shortlisted TSG model(s) 330 to generate a respective generated TSG data 345. In other words, the one or more shortlisted TSG model(s) 330 are tested / used for generating TSG data based on the user-provided TSG dataset 314.
[0051] The generated TSG data 345 may be synthetic data which retains the intrinsic statistical properties and temporal dependencies of the user-provided TSG dataset 314. In various embodiments, the generated TSG data 345 may be indicative on the suitability or relevance of the respective shortlisted TSG model(s) 330.
[0052] Further referring to FIG. 7, the benchmarking module 400 may further comprise an evaluator module 420 in signal communication between the first machine learning model 210and the user interface 902. The evaluator module 420 may be configured to perform one or more quantitative evaluations on the shortlisted TSG model(s) 330. The evaluator module 420 may also provide one or more evaluation measures as a suggestion / recommendation based on the shortlisted TSG model(s) 330. In various embodiments, the evaluator module 420 may perform one or more quantitative evaluations on each of the shortlisted TSG model(s) 330 based on one or more evaluation measures 375. The one or more evaluation measures 375 may be pre-stored in the TSG database 230. As such, the evaluator module 420 may obtain the one or more evaluation measures 375 from the TSG database 230 based on the user input 312 and / or the user-provided TSG dataset 314. As examples, the one or more evaluation measures 375 may comprise at least one of: an efficiency-based measure, a model-based measure, a feature-based measure, and a distance-based measure.
[0053] In various embodiments, a result of the one or more quantitative evaluations may act as a response 350 to the TSG query 320. In further embodiments as shown in FIG. 8, the evaluator module 420 may be configured to perform one or more quantitative evaluations on the shortlisted TSG model(s) 330 based on the respective generated TSG data 345. The TSG database 230 may comprise a plurality of evaluation measures 375. The plurality of evaluation measures 375 may comprise: efficiency-based measures, model-based measures, feature-based measures and distance-based measures.
[0054] Exemplary model-based measures may include but is not limited to: (i) Discriminative Score (DS) which employs a post-hoc time-series classification model with 2- layer GRUs or LSTMs to differentiate between original and generated series Each original series is labeled as real, while the generated series is labelled synthetic. Using these labels, an RNN classifier is trained. The classification error on a test set quantifies the generation model’s fidelity; (ii) Predictive Score (PS) which involves training a post-hoc time series prediction model on synthetic data. Using GRUs or LSTMs, the model predicts either the temporal vectorsof each input series for the upcoming steps or the entire vector. The model’s performance is then evaluated on the original dataset using the mean absolute error; (iii) Contextual-FID (C- FID) which extends the concept of Frechet Inception Distance (FID) from image generation to TSG. C-FID quantifies how well the synthetic time series conforms to the local context of the time series. Using the time series embeddings from, it learns embeddings that seamlessly blend with the local context.
[0055] Exemplary feature-based measures may include but is not limited to: (i) Marginal Distribution Difference (MDD) which computes an empirical histogram for each dimension and time step in the generated series, using the bin centers and widths from the original series. It then calculates the average absolute difference between this histogram and that of the original series across bins, assessing how closely the distributions of the original and generated series align; (ii) Auto Correlation Difference (ACD) which computes the autocorrelation of both the original and generated time series, then determines their differences. By contrasting the autocorrelations, we could evaluate how well dependencies are maintained in the generated time series; (iii) Skewness Difference; (iv) Kurtosis Difference (KD), (v) Training Time which refers to the wall clock time for training a TSG method, and is a vital measure for evaluating and deploying TSG methods due to economic considerations.
[0056] Exemplary distance-based measures may include but is not limited to: (i) Euclidean Distance (ED); and (ii) Dynamic Time Warping (DTW) which captures the optimal alignment between series regardless of their pace or timing. The alignment facilitated by DTW offers insights into the predictive quality of the generated series. Moreover, multi-dimensional DTW can enhance downstream classification tasks, serving as a discriminative measure.
[0057] In various embodiments, a result from the one or more quantitative evaluations may be provided and displayed on the user interface 902 as a response 350 to the user prompt 10.In various embodiments, a result of the one or more quantitative evaluations may be displayed on the user interface 902.
[0058] In various embodiments, the user-provided TSG dataset 314 and the respective generated TSG data 345 may be displayed on the user interface 902 using at least one visualization tool. Visualization offers an intuitive and visually interpretive perspective to directly compare and contrast the structures and patterns between the original and generated time series. In various embodiments, the at least one visualization tool may be selected from a group of visualization tools, such as: Principal Component Analysis (PCA); t-distributed Stochastic Neighbour Embedding (t-SNE) which allows succinctly visualizing the distribution of generated time series compared to the original one within a two-dimensional space; and Distribution Plot which illuminates the difference between the input and generated time series in terms of density, spread, and central tendency to show how the generated time series closely mirrors the original’s statistics.
[0059] In various embodiments, the system 100 may be provided with a data profiling module configured to provide insights relating to TSG dataset and TSG models, based on characteristics of the user-provided TSG dataset 314 and user requirements / needs derived from the user input 312. The system 100 may provide one or more insights for selecting the shortlisted TSG model(s) 330 as a response to the TSG query 320. The one or more insights may be an output of the data profiling module, with one or more insights pre-stored in the TSG database 230. The one or more insights may correspond to observations, indications and / or guidance relating to TSG model selection. As an example, the insights may include: data dimensionality, sequence length, and / or domain-specific application. The data profiling module may extract critical insights from multivariate time series datasets. These insights inform model selection, hyperparameter tuning, and evaluation metric interpretation, thereby enhancing both performance and interpretability.
[0060] Further referring to FIG. 8, the benchmarking module 400 may comprise a simulation module 410 in signal communication with an evaluator module 420. The benchmarking module 400 may be in signal communication between the first machine learning model 210 and the user interface 902. Hence, the simulation module 410 may be configured to simulate one or more generated TSG data 345 based on the one or more shortlisted TSG model(s) 330. The one or more generated TSG data 345 may be provided to the evaluator module 420 to perform one or more quantitative evaluations on the one or more generated TSG data 345 which is indicative of the performance of the shortlisted TSG model(s) 330.
[0061] According to another aspect of the current disclosure, FIG. 9 illustrates a method 500 of Time Series Generation (TSG) model selection. According to various embodiments, the method 500 comprises: in 510, receiving a user prompt, theuser prompt comprising a user input and a user-provided TSG dataset; in 520, providing the user input as input to a first machine learning model, to detennine the user prompt as a TSG query; in 530, providing the user input and the user-provided TSG dataset as input to a second machine learning model, to select at least one shortlisted TSG model from a TSG database; and in 540, using the first machine learning model, providing the at least one shortlisted TSG model as a response to the TSG query.
[0062] The method 500 may be iterative 505 in nature. Hence, the method 500 may further include: responsive to a further user input, providing at least one subsequent shortlisted TSG model based on the further user input and the user-provided TSG dataset, wherein the at least one subsequent shortlisted TSG model is non-identical to the at least one shortlisted TSG model.
[0063] In various embodiments, the method 500 further comprises: in 550, preprocessing the user-provided TSG dataset to obtain a preprocessed TSG dataset, and determining a context of the TSG query based on the preprocessed TSG dataset. In various embodiments, the method500 further comprises: in 560, providing the context of the TSG query and the user-provided TSG dataset as input to the second machine learning model, to select the at least one shortlisted TSG model from the TSG database.
[0064] In various embodiments, the method 500 further comprises: in 570, providing the user-provided TSG dataset to each of the at least one shortlisted TSG model; and generating a respective generated TSG data corresponding to a respective one of the at least one shortlisted TSG model. Alternatively, the method 500 further comprises: in 580, providing a pre-stored TSG dataset from the TSG database to each of the at least one shortlisted TSG model; and generating a respective generated TSG data corresponding to a respective one of the at least one shortlisted TSG model.
[0065] In various embodiments, the method 500 further comprises: in 590, based on at least one evaluation measure, performing a quantitative evaluation on each of the at least one shortlisted TSG model based on the respective generated TSG data. In various embodiments, the method 500 further comprises: providing at least one evaluation measure based on each of the at least one shortlisted TSG model.
[0066] Exemplary Embodiment of proposed system and method - TSGAssist
[0067] FIG. 10 illustrates an exemplary embodiment of the proposed system configured as an interactive assistant (TSGAssist), offering an intuitive, user-friendly interface for industry professionals. TSGAssist comprises two key components: TSG Recommender and TSG Benchmarking module.
[0068] TSG Recommender
[0069] TSG Recommender comprises and utilizes Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) to provide personalized, context-aware recommendations for TSG methods and evaluation metrics. TSG Recommender uses conversational interface bridges the knowledge gap, allowing users to interact naturally withthe system and receive custom advice based on their specific industry needs and data characteristics. TSG Recommender is configured to provide personalized suggestions and information This module allows users to engage with TSGAssist through a chatbot-style interface. As illustrated in FIG. 11, the module operates through a series of interconnected components, employing RAG and LLMs to harness insights from a Knowledge Database and provide user-specific recommendations.
[0070] A user begins by interacting with the system, entering TSG-related queries, and uploading datasets for customized advice. Upon receiving the user’s input, this component preprocesses the time series data to better grasp the context and specifics of each request. Concurrently, the knowledge base is accessed, which includes detailed information on TSG methods, metrics, and datasets essential for up-to-date and relevant recommendations. At the core of TSG Recommender lies the Retrieval component, driven by the principles of RAG and LLMs and configured to mine or retrieve information from the knowledge base, in response to user queries. After retrieval, the LLM component processes the user’s prompt and fetched data, creating responses that are informative, coherent, and context-appropriate. The LLM’s natural language understanding and text generation aids in enabling a seamless and engaging user experience. The system may then select and deliver recommendations (such as shortlisted TSG models) and / or information to the user.
[0071] Retrieval-Augmented Generation (RAG) Model
[0072] The RAG model may be implemented to support hybrid reasoning over LLMs augmented with user-uploaded knowledge files. For example, the RAG model may be implemented via the OpenAI Assistants. A “Recommender class” is defined and serves as the interface between user queries and the assistant's response generation pipeline. The initialization sequence involves four steps: (1) Configuration of System Prompt: A domainspecific prompt that constrains the assistant to focus on time series generation, model selection, 1and benchmarking. (2) Knowledge Base Incorporation: A curated knowledge file, typically a structured summary of TSG methods, datasets, metrics, and use-case mappings, is uploaded and indexed for retrieval (3) Assistant Configuration: The assistant is instantiated with retrieval tools enabled and linked to the uploaded file. (4) Thread Context: A session-level thread is created to maintain conversational memory across turns.
[0073] When the user submits a query, TSGAssist dynamically retrieves relevant portions of the knowledge base and integrates them into its generation pipeline. The response is then synthesized with embedded citations stripped for cleaner output. To aid in robustness, the assistant's response is polled at fixed intervals until either a response is generated or a timeout threshold is reached. This looped polling ensures that long-running queries (e g., complex retrievals) are handled gracefully Errors such as expired or failed sessions are detected early, and appropriate fallbacks are triggered.
[0074] The integration of RAG brings several practical benefits such as: (i) Personalization: Users receive tailored recommendations that consider their domain and dataset characteristics, (ii) Explainability: TSG Models can be described and explained contextually, including their strengths, assumptions, and limitations, (iii) Task Relevance: RAG is able to distinguish between different downstream tasks from the descriptions of the users such as forecasting, imputation, or simulation, and suggest suitable metrics or data augmentations, (iv) Accessibility: Industrial users unfamiliar with deep generative modelling may access insights and best practices without the requiring the laborious process of reading extensive documentation or source code
[0075] The TSG Recommender uniquely supports continuous dialogue, letting users iteratively adjust their queries in the same conversation based on initial suggestions. This capability fosters in-depth exploration and clarification, thereby improving the decision-making process with more precise and tailored recommendations.
[0076] Data profiling module
[0077] Effective TSG hinges not only on the architecture of the generative model but also on the ability to tailor modelling strategies to the intrinsic characteristics of the input data. A systematic data profiling module may be provided configured to extract insights from multivariate time series datasets. These insights inform model selection, hyperparameter tuning, and evaluation metric interpretation, thereby enhancing both performance and interpretability.
[0078] Time series data often exhibit substantial variability across domains in terms of temporal length, feature dimensionality, series heterogeneity, periodic structure, and noise complexity. A one-size-fits-all generative approach may not be optical. The proposed system may perform dataset characterization along four dimensions (i.e., size, series diversity, periodicity, and structural complexity). These descriptors of the four dimensions may support downstream tasks such as (1) selecting appropriate generative models, (2) constructing datasetspecific prompts for conditional generation or interpretation, and (3) enabling fair and contextualized benchmarking across datasets.
[0079] For example, given a 3D input tensor X E ]RJVX7’XCrepresenting N sequences of length T with C channels, the following indicators were extracted: (i) Size: Classified as small or large based on a threshold of N = 1000 sequences. This affects the statistical power of training and evaluation metrics, (ii) Series Diversity: A dataset is deemed homogeneous if the average pairwise Pearson correlation across sensor channels exceeds 0.5. Otherwise, it is considered heterogeneous, indicating varying signal dynamics across channels, (iii) Periodicity: Periodicity is estimated by computing the autocorrelation function (ACF) for each channel across a subset of sequences and averaging the values. A dataset is flagged as periodic if any lag-fc (k>0) ACF exceeds a predefined threshold (0.5), revealing the presence of seasonal cycles or repeating structures, (iv) Structural Complexity: Datasets with more than 10 channels( C > 10 ) are considered complex, requiring models that can handle high-dimensional interdependencies.
[0080] Prompt construction module
[0081] To bridge the gap between dataset profiling and practical modelling choices, an automated prompt construction module is provided. Based on the extracted metadata, a structured natural language prompt is assembled to guide the users The prompt provides a concise and human-readable summary of the dataset's nature, while hiding low-level technical tenns that may overwhelm non-experts. This mechanism also serves as the foundation for building interactive chat-based interfaces for TSGAssist.
[0082] Knowledge mining module
[0083] The proposed knowledge mining module not only enhances transparency in generative modeling but also reduces manual effort in dataset curation and experimental diagnostics. For example, high series diversity may suggest the suitability of attention-based architectures, while strong periodicity might favour models with temporal filters or Fourier transforms. Furthermore, the analysis pipeline helps detect unbalanced benchmark settings where model comparisons may be misleading due to overlooked data characteristics. Overall, this layer of data introspection enables more principled, adaptive, and explainable use of generative models in time series domains.
[0084] TSG Benchmarking module
[0085] TSG Benchmarking module (or the TSG Benchmarking platform) enables users to apply shortlisted or recommended TSG methods to the user-provided TSG data directly on the platform. It supports instant comparison and evaluation with a wide range of metrics, providing practical insights into the application of various methods on real datasets.
[0086] Guided by TSG Recommender, the TSG Benchmarking module allows users to assess the suggested TSG methods and measures tailored to their specific tasks. This moduleenables users to apply these methods to the user-provided datasets and / or investigate results using sample datasets, enhancing their understanding.
[0087] The TSG Benchmarking module has three components: (i) Initialization: Users begin by choosing TSG methods and evaluation metrics from the control panel, then move to the dedicated benchmarking platform. This selection is guided by TSG Recommender’s detailed advice, ensuring evaluations are targeted and pertinent, (ii) Quantitative Evaluations: It includes comprehensive metrics across efficiency -based, model-based, feature-based, and distancebased categories from the Knowledge Base. The efficiency-based metrics may track each method’s running time, offering insights into efficiency. This broad evaluation spectrum allows users to thoroughly assess TSG methods’ effectiveness from various angles, matching their specific requirements The chosen metrics are clearly displayed on the main interface, with options for detailed analysis of each, (iii) Visualizations: This component provides visualization tools like PCA, t-SNE, and Distribution Plot on the secondary tier for comparing original and generated time series. They are crucial for understanding how well TSG methods preserve core patterns and traits of time series data.
[0088] The TSG Benchmarking module is a dynamic update system Selecting a new TSG method automatically refreshes the evaluation results and visualizations to include it, keeping the module up-to-date and aligned with the latest developments.
[0089] Knowledge Base
[0090] At the core of TSGAssist's RAG component lies a well -structured and curated knowledge base that encodes essential information about TSG models, evaluation metrics, datasets, and best practices. The Knowledge Base provides a benchmarking framework for TSG methods The Knowledge Base comprises three core parts: (i) Data Preprocessing: A standardized pipeline for preprocessing real -world time series data, including normalization and segmentation, to ensure uniformity and comparability across different TSG methods, (ii)Methods and Insights: the Knowledge Base offers insights and guidance for selecting TSG methods based on data characteristics and needs. It highlights the significance of factors like data dimensionality, sequence length, and application domain, (iii) Evaluation Measures: the Knowledge Base stores / provides a suite of evaluation metrics, including model-based, featurebased, and distance-based measures. These metrics provide a comprehensive overview of each method’s performance
[0091] The effectiveness of a RAG-enhanced assistant is dependent on the quality and comprehensiveness of the corresponding knowledge base. In the domain of TSG, where methods are rapidly evolving and evaluation criteria are multi-faceted, a centralized knowledge repository helps bridge the cognitive gap between academic research and practical applications. The knowledge base provides factual grounding thus enabling the LLM's responses to be grounded in validated knowledge, reducing hallucinations; task-aware retrieval which supports task-specific selections or recommendations by encoding model capabilities, dataset characteristics, and metric suitability; and interactive exploration which allows users to interactively query about TSG methods and receive domain-informed, context-specific responses
[0092] The knowledge base provides systematic benchmarking and methodologies for TSG. Its contents are compiled into structured files and indexed for retrieval. The knowledge base comprises: (i) TSG Methods: Each TSG model is documented with the model type, core mechanisms, strengths and limitations, and the use case alignment, (ii) Evaluation Measures: The descriptions of quantitative and qualitative evaluation metrics across four dimensions (i.e., model-based, feature-based, distance-based, and visualization). Each metric is explained with its formulation, interpretability, and sensitivity to different types of generation errors, (iii) Metadata for Selected Dataset: The knowledge base contains summaries of benchmark datasetsincluding their dimensionality, sequence length, periodicity, and domain relevance. This information is used to match suitable models and evaluation protocols to each dataset.
[0093] The knowledge base is then serialized as a local file and uploaded during the system initialization. Once uploaded, it becomes retrievable via RAG-enabled LLM prompts. The assistant's instructions define the expected usage and boundaries of this knowledge, ensuring that all generated responses remain faithful to the structured content. It not only acts as a static information source for retrieval, but also a dynamic decision engine when combined with user queries and dataset profiling results, enabling delivery of contextualized TSG model selections or recommendations.
[0094] Implementation of TSGAssist
[0095] In an exemplary implementation, TSGAssist is implemented as a standalone web application developed with Python 3.7 and the Dash framework, providing a user-friendly GUI. FIGs. 12A and 12B show exemplary screenshots of TSGAssist, offering insight into its functionality. TSGAssist includes basic functions like a control panel for data upload, TSG method and measure selection (FIG. 12A (a.l)), along with a Time Series Overview canvas for data selection and raw MTS visualization (FIG. 12A (a 2)). There are two principal modules in TSGAssist that enhance user interaction: (1) TSG Recommender for personalized method suggestions, and (2) TSG Benchmarking for hands-on evaluation and comparison of TSG methods.Table 1. Exemplary Pseudo Code
[0096] Table 1 illustrates a pseudo code for a workflow in TSGAssist. The workflow begins when a user interacts with TSGAssist, the user begins by uploading a time series dataset and entering a textual query (or user input) through the web interface. The system handles this interaction by capturing the uploaded dataset and the user's textual query. The uploaded data is decoded and stored using the Datastorage class, which extracts the structural attributes of the data such as its size, series dimensionality, and length. To further contextualize the user's textual query, the system uses the examine data featureQ function to analyze the data's properties such as specifically, its size, diversity, periodicity, and complexity.
[0097] These extracted characteristics (or context) are used to construct a contextual prompt via the assemble_prompt() function, and is appended to the user's prompt. The purpose is to inform the LLM of the data context without explicitly disclosing it in the assistant's reply. For instance, the constructed prompt may start with: “The dataset I am to handle is large, heterogeneous, periodic, complex...” followed by instruction not to reveal this information directly. This augmented prompt is then passed into the response generation module.
[0098] Once the user prompt is assembled, it is submitted to the Recommender class, which interfaces with the LLM. If initialized, the Recommender sets up an LLM and uses RAG byincorporating the knowledge base file. It is also linked to a persistent thread and configured to use the uploaded knowledge for contextually accurate and grounded replies.
[0099] When the user submits a new user input (textual input), the assistant receives it as a new message in the conversation thread. A retrieval -enhanced generation run is initiated, during which the assistant searches the knowledge base for relevant content. This retrieved content is used to inform the response generated by the LLM. The system polls the assistant's run status in a loop until completion, after which it retrieves the assistant's reply. The annotations from the retrieval step are removed to ensure a clean user-facing message.
[0100] The final response i s returned to the web interface and inserted into the chat window, replacing the placeholder message (e.g. “Thinking”) that was initially displayed. The user can continue the conversation in an iterative manner, and the system retains contextual continuity within the same chat thread. This conversational loop allows for progressively refined and personalized TSG recommendations, making use of both the user's inputs and the characteristics of the uploaded data.
[0101] The benchmarking module in TSGAssist allows users to apply the shortlisted (or recommended) TSG methods and evaluation metrics directly on their uploaded data This module operates through three tightly integrated components. First, users initialize the benchmarking session by selecting TSG methods and evaluation measures via the control panel. Next, quantitative evaluations are performed using metrics sourced from TSGBench, which include model-based, feature-based, and distance-based criteria. These evaluations are presented with corresponding error bars to convey statistical robustness.
[0102] The system also provides visual benchmarking through dimensionality reduction by t-SNE and distribution plots. Users may visually compare the original and synthetic data distributions across different shortlisted TSG methods. The dynamic nature of the system allows visualizations and metric displays to be updated in real time as users change the user-provided TSG dataset and / or TSG methods. The integrated configuration supports informed decision-making by offering both analytical and visual evidence of TSG performance.
[0103] Experiments on TSG Assist
[0104] TSGAssist is effective and easy to use through various scenarios, such as: (1) enhancing understanding and application of TSG in real-world contexts through RAG, (2) offering tailored recommendations for specific industry requirements, and (3) providing a comprehensive benchmarking platform for users to assess and select the most appropriate TSG methods and metrics for their applications. These scenarios demonstrate how TSGAssist effectively mitigates the challenges faced by industry professionals in navigating the complex landscape of TSG, fostering greater application of TSG in various industrial domains.
[0105] Experiments were performed on the listed scenarios, each showing a distinct facet of TSGAssist’ s capabilities. The primary goals are: (1) showcasing the efficacy of integrating RAG in the TSG Recommender, (2) offering domain-specific recommendations for TSG methods and evaluation measures, and (3) facilitating thorough evaluation of various TSG methods.
[0106] Scenario 1 : RAG Efficacy
[0107] The first scenario establishes the foundational effectiveness of TSG Recommender’s RAG component.
[0108] Objective: The goal is to show the improvement in the relevance and faithfulness of responses from an LLM when augmented with RAG, as opposed to those by the LLM alone.
[0109] Setup: 10 participants were engaged. The participants were given a series of prompts to input into two systems: a standalone LLM and a RAG-enhanced LLM in TSGAssist. The responses from both systems were displayed side by side anonymously, with participants unaware of each system’s contribution.
[0110] Questions: (1) Relevance: Participants assessed which responses are more relevant to their initial prompt, gauging the system’s ability to effectively understand and address the query. (2) Faithfulness: Participants were asked to evaluate which of the two responses appears to be more consistent with existing knowledge in TSGBench, thereby assessing the faithfulness of the information provided.[001 11] According to the results presented in FIG 1 A, more than 70% of users found the RAG-enhanced LLM in TSG Recommender to be significantly more effective in both relevance and faithfulness compared to the standalone LLM. The RAG-enhanced system’s responses were more aligned with the knowledge base, indicating a higher degree of factual accuracy compared to the standalone LLM. Responses from TSGAssist were consistently rated as more pertinent to user prompts, demonstrating a better understanding of user queries. These results affirm the efficacy of integrating RAG with LLMs, leading to improvements in generating accurate and relevant text, which is crucial for TSG Recommender.
[0112] Scenario 2: Domain-Specific Applications
[0113] Leveraging the demonstrated success of the TSG Recommender’s RAG component in Scenario 1, Scenario 2 shifts focus towards its performance in domain-specific contexts, underscoring its adaptability and precision in providing tailored recommendations for TSG methods and evaluation measures.
[0114] Objective: The aim is to showcase TSGAssist’s capability to generate recommendations that are not only methodologically sound but also contextually relevant to specific domains.
[0115] Setup: Participants from finance, environmental science, and manufacturing are provided with responses generated by TSG Recommender tailored to each domain.
[0116] Questions: For each question, participants were instructed to rate the recommendations using a 1 to 5 scale based on either (1) Method Comprehensiveness or (2)Measure Appropriateness, along with (3) Overall Domain Helpfulness. In particular, (1) Method Comprehensiveness: Participants evaluated the applicability of recommended TSG methods in their field. (2) Measure Appropriateness: Participants rated the relevance of TSG evaluation measures for their domain. (3) Overall Domain Helpfulness: Participants scored the overall effectiveness of TSGAssist in meeting their specific domain needs.[001 17] The average rating for all questions across various domains consistently exceeded 4.2, as shown in FIG. 13B. The TSG methods and evaluation measures recommended by TSGAssist were well -regarded for their relevance and practicality in specific domains. Overall, the responses affirmed TSGAssist’ s proficiency in effectively addressing domain-specific queries.[001 18] Scenario 3: TSG Benchmarking Module / Platform
[0119] The third scenario presents TSGAssist as a holistic TSG benchmarking platform suitable for academic and industry professionals. It aids in evaluating TSG methods tailored to specific needs. For example, a manufacturing sector user uploads sensor-based time series into TSGAssist (Figure 12A (a.l)), seeking the best TSG methods for their data traits and predictive maintenance needs.
[0120] Recommendations: TSG Recommender (FIG. 12A (a.3)) provided with VAE-based methods like TimeVAE and TimeVQVAE for TSG It also suggests COSCI-GAN for its proficiency in handling complex multivariate relationships, crucial in machinery interactions. For evaluation, TSGAssist selected and recommended using Predictive Score (PS) to compare models trained on synthetic and original data for maintenance predictions It also endorses C- FID and ACD to verify the preservation of temporal correlations in synthetic data, reflecting real machinery operation patterns.
[0121] Benchmarking: Users begin on the benchmarking platform by selecting recommended methods and metrics from the control panel (FIG. 12B (b.l)). Comparativeresults (FIG. 12B (b.2)) highlight TimeVAE and C0SC1-GAN as top choices, noting TimeVAE’s significantly faster training compared to GAN-based alternatives. Additional methods like TimeGAN can be included for a more thorough analysis. Users can compare synthetic and original data by PCA, t-SNE, or Distribution Plot (FIG. 12B (b.3)), further confirming the suitability and effectiveness of the chosen methods.
[0122] User Experience and Impact
[0123] Industry users (such as manufacturing managers and professionals) are able to seamlessly integrate recommended TSG methods and evaluation metrics into their workflows using TSGAssist without delving into extensive TSG literature. Its accessibility and practicality significantly simplify the adoption and integration of advanced TSG solutions into real-world applications, improving predictive maintenance strategies and operational efficiency.
[0124] The proposed system and method may be implemented by a processor system 900 as illustrated in the schematic block diagram of FIG. 14. Components of the processing system 900 may be provided within one or more computing device to carry out the functions of the modules or any other modules. One skilled in the art will recognize that the exact configuration or arrangement illustrated in FIG. 14 is provided by way of example only, e g., each processing system provided may be different and the exact configuration of processing system 900 may vary.
[0125] In embodiments of the present disclosure, the processing system 900 may include a controller 901 and user interface 902. User interface 902 is configured to enable manual interactions between a user and the computing module as required. For this purpose, the processing system 900 includes the input / output components required for the user to enter instructions to provide updates to each of the modules. A person skilled in the art will recognize that components of user interface 902 may vary from embodiment to embodiment but may typically include one or more input devices 935 such as but not limited to a touchscreen, akeyboard, a joystick, a mouse, a microphone, etc. The user interface 902 can also include a media player 940, which can be in the form of one or more playback devices, including but not limited to a display, a speaker, earphones, headsets, etc.
[0126] The controller 901 is configured to be in data communication with the user interface 902 via bus 915. The controller 901 includes memory 920 and processor 905 mounted on a circuit board to process instructions and data, e.g., to perform the method of the present disclosure. The controller 901 includes an operating system 906, an input / output (I / O) interface 930 for communicating with user interface 902, and a communications interface, e.g., a network card 950. The network card 950 may, for example, be configured to send data from the controller 901 via a wired or wireless network to other processing devices or to receive data via the wired or wireless network. Wireless networks that may be utilized by the network card 950 include, but are not limited to, Wireless-Fidelity (Wi-Fi), Bluetooth, Near Field Communication (NFC), cellular networks, satellite networks, telecommunication networks, Wide Area Networks (WAN), and etc.
[0127] Memory 920 and operating system 906 are in data communication with central processing unit (CPU) 905 via bus. The memory 920 may include both volatile and non-volatile memory. The memory 920 may include more than one of each type of memory, e.g., Random Access Memory (RAM) 923, Read Only Memory (ROM) 925, and a mass storage device 927. The mass storage device 927 may include one or more solid-state drives (SSDs). One skilled in the art will recognize that the memory described above includes non-transitory computer- readable media and shall be taken to include all computer-readable media except for a transitory, propagating signal. Typically, instructions are stored as program code in the memory but can also be hardwired. Memory 920 may include a kernel and / or programming modules such as a software application that may be stored in either volatile or non-volatile memory.
[0128] Herein, the term “processor” is used to refer generically to any device or component that can process computer-readable instructions, including for example, a microprocessor, microcontroller, programmable logic device, or other computational device. That is, processor 905 may be provided by any suitable logic circuitry for receiving inputs, processing them in accordance with instructions stored in memory, and generating outputs (for example to the memory components or media player 940). In the present disclosure, processor 905 may be a single core or multi-core processor with memory addressable space. In one example, processor 905 may be multi-core, comprising — for example — an 8 core CPU. In another example, it could be a cluster of CPU cores operating in parallel to accelerate computations.
[0129] Further, one skilled in the art will recognize that certain functional units in this description have been labelled as modules throughout the specification. The person skilled in the art will also recognize that a module may be implemented as circuits, logic chips or any sort of discrete component. Still further, one skilled in the art will also recognize that a module may be implemented in software which may then be executed by a variety of processor architectures. In embodiments of the disclosure, a module may also comprise computer instructions or executable code that may instruct a computer processor to carry out a sequence of events based on instructions received. In further embodiments, the module may comprise a combination of different types of modules or sub-modules. The choice of the implementation of the modules may be determined by a person skilled in the art and does not limit the scope of the claimed subject matter in any way.
[0130] All examples described herein, whether of apparatus, methods, materials, or products, are presented for the purpose of illustration and to aid understanding, and are not intended to be limiting or exhaustive. Modifications may be made by one of ordinary skill in the art without departing from the scope of the invention as claimed.
Claims
CLAIMS1. A system, comprising: memory storing instructions; and a processor coupled to the memory and configured to process the stored instructions to implement: a module configured to perform a method of Time Series Generation (TSG) model selection, the method including: receiving a user prompt, the user prompt comprising a user input and a user- provided TSG dataset; providing the user input as input to a first machine learning model, to determine the user prompt as a TSG query; providing the user input and the user-provided TSG dataset as input to a second machine learning model, to select at least one shortlisted TSG model from a TSG database; and using the first machine learning model, providing the at least one shortlisted TSG model as a response to the TSG query.
2. The system as recited in claim 1, wherein the method further comprising: preprocessing the user-provided TSG dataset to obtain a preprocessed TSG dataset, and determining a context of the TSG query based on the preprocessed TSG dataset.
3. The system as recited in claim 2, wherein the method further comprising: providing the context of the TSG query and the user-provided TSG dataset as input to the second machine learning model, to select the at least one shortlisted TSG model from the TSG database.
4. The system as recited in any one of claims 2 and 3, wherein the context of the TSG query comprises a domain-specific application of the user-provided TSG dataset.
5. The system as recited in any one of claims 2 to 4, wherein preprocessing the user- provided TSG dataset comprises: normalizing and segmenting the user-provided TSG dataset.
6. The system as recited in any one of claims 1 to 5, wherein the method further comprising: providing the user-provided TSG dataset to each of the at least one shortlisted TSG model; and generating a respective generated TSG data corresponding to a respective one of the at least one shortlisted TSG model.
7. The system as recited in any one of claims 1 to 5, wherein the method further comprising: providing a pre-stored TSG dataset from the TSG database to each of the at least one shortlisted TSG model; and generating a respective generated TSG data corresponding to a respective one of the at least one shortlisted TSG model.
8. The system as recited in any one of claims 6 and 7, wherein the method further comprising: based on at least one evaluation measure, performing a quantitative evaluation on each of the at least one shortlisted TSG model based on the respective generated TSG data.
9. The system as recited in claim 8, wherein the at least one evaluation measure comprising at least one of: an efficiency-based measure, a model-based measure, a featurebased measure, and a distance-based measure.
10. The system as recited in any one of claims 8 and 9, wherein the method further comprising: displaying a result of the quantitative evaluation.
11. The system as recited in claim 10, wherein displaying the result of the quantitative evaluation comprising: displaying the user-provided TSG dataset and the respective generated TSG data using at least one visualization tool.
12. The system as recited in claim 11, wherein the at least one visualization tool is selected from: Principal Component Analysis (PCA), t-distributed Stochastic Neighbour Embedding (t- SNE), and Distribution Plot.
13. The system as recited in any one of the above claims, wherein the method further comprising: providing at least one insight for selecting the at least one shortlisted TSG model as the response to the TSG query, wherein the at least one insight comprising at least one of: data dimensionality, sequence length, and domain-specific application.
14. The system as recited in any one of the above claims, wherein the first machine learning model is a large language model, and the second machine learning model is a retrieval- augmented generation model.
15. The system as recited in any one of the above claims, wherein the TSG database comprises a plurality of TSG models, a plurality of evaluation measures and a plurality of prestored TSG datasets.
16. The system as recited in any one of the above claims, wherein the user input comprises a series of natural language input from a user to the first machine learning model.
17. The system as recited in any one of the above claims, wherein the method further comprising: responsive to a further user input, providing at least one subsequent shortlisted TSG model based on the further user input and the user-provided TSG dataset, wherein the at least one subsequent shortlisted TSG model is non-identical to the at least one shortlisted TSG model.
18. The system as recited in any one of the above claims, wherein the method further comprising: providing at least one evaluation measure based on each of the at least one shortlisted TSG model.
19. The system as recited in any one of the above claims, wherein the method further comprising: receiving the user prompt via a user interface, and providing the at least one shortlisted TSG model via the user interface.
20. A portable device, comprising the system as recited in any one of the above claims.
21. The portable device as recited in claim 20, further comprising a user interface and an input device.
Citation Information
Patent Citations
Method and device for predicting and enhancing query instruction of electric power customer service system
CN118152428A
Interface for Visualizing and Improving Model Performance
US20200118018A1
Prompt generator for use with one or more machine learning processes
US20240095077A1
System and method for efficient language model editing using contextual prompt generator
US20240104309A1
Automatic forecasting using meta-learning
US20240152769A1
Cited By
Industrial data intelligent query method and system based on large language model
CN121833807A