Computer-implemented method, computer system and computer program (basic model for a dynamic system)

The method uses a dynamic system dictionary and diffusion decoder to generate training data for dynamic systems, addressing the challenge of obtaining sufficient labeled data for predictive models, enhancing model training efficiency.

JP2026031396APending Publication Date: 2026-02-24INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025091389
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-07
Filing Date
2025-05-30
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Measuring and obtaining extended time series data for dynamic physical systems is difficult and expensive, and existing methods struggle with generating sufficient labeled training data for predictive models.

Method used

A computer-implemented method involving a dynamic system dictionary, classification through constraint learning, and a diffusion decoder to generate time-series segments, utilizing an encoder and noise generator to create training data for dynamic systems.

Benefits of technology

Generates new time series data that mimic the dynamics of input samples, providing ample training data for predictive models, thus improving the efficiency and effectiveness of model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026031396000001_ABST
    Figure 2026031396000001_ABST
Patent Text Reader

Abstract

Physical systems are often dynamic and constantly changing. Measuring these physical systems can be a difficult and expensive task.SOLUTION: Providing a plurality of dynamic systems to a dynamic systems dictionary, wherein the dynamic systems dictionary may comprise a library of functions, classifying each of the plurality of dynamic systems, wherein the classifying may comprise generating hierarchical dynamic systems data based on constraint learning, including an encoder and a noise generator, training a diffusion decoder to generate time series segments based on the classified plurality of dynamic systems, providing a first time series data segment, and using the first time series data segment as input; An approach is provided that can include generating time series dynamic training data based on a diffusion decoder.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to artificial intelligence and machine learning models, and more particularly to developing foundational models for dynamical systems. [Background technology]

[0002] Physical systems are often dynamic and constantly changing. Measuring these physical systems can be a difficult and expensive task. Furthermore, obtaining extended time series data for downstream prediction tasks is complex and difficult to obtain. It would be advantageous to identify a similar dynamical system that can best explain the time series data associated with a physical system and use the identified dynamical system to learn the behavior of that physical system.

[0003] Base models are developed in the following general manner: First, data related to the phenomenon of interest is collected at a large scale. A model is trained using a portion of the data, and if the model correctly predicts the input, a reward is given. The model is evaluated in this manner until satisfactory results are obtained. The model can then be used as a base model from which the pre-trained model is developed for more specific downstream applications related to the initial phenomenon of interest. Summary of the Invention [Problem to be solved by the invention]

[0004] Physical systems are often dynamic and constantly changing, and measuring these physical systems can be a difficult and expensive task. [Means for solving the problem]

[0005] According to one embodiment of the present invention, there is provided a computer-implemented method for generating time-series dynamic system training data. The computer-implemented method includes providing a plurality of dynamic systems to a dynamic system dictionary, where the dynamic system dictionary includes a library of functions. The computer-implemented method may further include classifying each of the plurality of dynamic systems, where the classifying includes generating hierarchical dynamic system data based on constraint learning including an encoder and a noise generator. The computer-implemented method further includes training a diffusion decoder to generate time-series segments based on the classified plurality of dynamic systems. Furthermore, the computer-implemented method may include providing a first time-series data segment; and generating time-series dynamic training data based on the diffusion decoder using the first time-series data segment as an input.

[0006] According to another embodiment of the present invention, there is provided a computer system for generating time-series dynamic system training data. The system includes a memory and a processor in communication with the memory, wherein the processor is configured to perform one or more operations. The operations may include providing a plurality of dynamic systems to a dynamic system dictionary, wherein the dynamic system dictionary includes a library of functions. The operations may further include classifying each of the plurality of dynamic systems, wherein the classifying step includes generating hierarchical dynamic system data based on constraint learning, the classifying step including an encoder and a noise generator. The operations may also include training a diffusion decoder to generate time-series segments based on the classified plurality of dynamic systems. Furthermore, the operations may include providing a first time-series data segment and generating time-series dynamic training data based on the diffusion decoder using the first time-series data segment as an input.

[0007] Another embodiment of the present invention may include a computer program product for generating time-series dynamic system training data. The computer program product may include a computer storage device and program instructions stored on the computer storage device. The program instructions may include program instructions for providing a plurality of dynamic systems to a dynamic system dictionary, where the dynamic system dictionary includes a library of functions. The program instructions may also include program instructions for classifying each of the plurality of dynamic systems, where the classifying step includes generating hierarchical dynamic system data based on constraint learning including an encoder and a noise generator. The embodiment may also include program instructions for training a diffusion decoder to generate time-series segments based on the classified plurality of dynamic systems. The embodiment may further include program instructions for providing a first time-series data segment and for generating time-series dynamic training data based on the diffusion decoder using the first time-series data segment as an input. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a block diagram illustrating an exemplary computing environment according to one embodiment of the present invention.

[0009] [Figure 2] Figure 2A is a block diagram illustrating an exemplary system for generating dynamic system training data, according to one embodiment of the present invention. Figure 2B is a block diagram illustrating a dynamic system training data generation engine 200, according to one embodiment of the present invention.

[0010] [Figure 3] FIG. 3 is a block diagram illustrating steps for generating time series training data with a function dictionary 300 according to one embodiment of the present invention.

[0011] [Figure 4] 1 is a flowchart illustrating steps for generating dynamic system training data for a given time series according to one embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0012] Embodiments of the present invention recognize the advantages associated with a system or methodology for dynamic system foundation models ("DS-FM"). Labeled training data for dynamic systems is often scarce. Embodiments of the present invention recognize the need to develop or generate labeled training data from actual samples of dynamic systems that have similar dynamics but are not perfectly similar, which can lead to overfitting. DS-FM learns the dynamics of time series data for a dynamic system from small segments of sample data. Embodiments of the present invention may utilize a dynamic system dictionary for various dynamic systems. Essentially, an AI or machine learning model can utilize the dynamic system dictionary to recognize patterns in segments of sample data, predict behavior, and generate new samples for training a dynamic system predictive model.

[0013] Embodiments of the present invention improve upon current techniques in that they can generate new time series data that follow dynamics substantially similar to that of a given input sample, thereby ensuring substantially more time series data or dynamic function data that can be used to train a predictive model.

[0014] As the embodiments are described in detail with reference to the figures, it should be noted that references herein to "one embodiment," "another embodiment," etc. indicate that the described embodiment may include a particular feature, structure, or characteristic, but not all embodiments may necessarily include that particular feature, structure, or characteristic. Moreover, such phrases do not necessarily refer to the same embodiment. Furthermore, by describing a particular feature, structure, or characteristic in the context of one embodiment, one skilled in the art will have knowledge that such feature, structure, or characteristic also affects other embodiments, whether or not explicitly described.

[0015] Various aspects of the present disclosure are illustrated by text, flowcharts, block diagrams of computer systems, and / or block diagrams of machine logic included in computer program product (CPP) embodiments. For any flowchart, depending on the technology involved, operations may be performed in an order different from that shown in a given flowchart. For example, depending again on the technology involved, two operations shown in successive flowchart blocks may be performed in reverse order, as a single integrated step, simultaneously, or in an at least partially overlapping manner.

[0016] A computer program product embodiment ("CPP embodiment" or "CPP") is a term used in this disclosure to describe any set of one or more storage media (also referred to as "media"), collectively contained in one or more storage devices, that collectively contain machine-readable code corresponding to instructions and / or data for performing the computer operations specified in a given CPP claim. A "storage device" is any tangible device that can hold and store instructions for use by a computer processor. The computer-readable storage medium may be, but is not limited to, an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, a mechanical storage medium, or any suitable combination of the foregoing. Some known types of storage devices that include these media include diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices (such as punch cards or pits / lands formed on a major surface of a disk), or any suitable combination of the foregoing. Computer-readable storage media, as the term is used in this disclosure, is not to be construed as storage in the form of transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide, light pulses passing through fiber optic cables, electrical signals communicated through wires, and / or other transmission media.As will be appreciated by those skilled in the art, data is typically moved at some infrequent time during the normal operation of a storage device, such as during access, defragmentation, or garbage collection, but this does not make the storage device temporary because the data is not temporary while it is stored.

[0017] Reference is now made to Figure 1, which is a block diagram illustrating an exemplary computing environment according to one embodiment of the present invention. In addition to a volumetric object adaptation engine 200, the computing environment 100 includes, for example, a computer 101, a wide area network (WAN) 102, an end user device (EUD) 103, a remote server 104, a public cloud 105, and a private cloud 106. In this embodiment, the computer 101 includes a processor set 110 (including processing circuitry 120 and cache 121), a communications fabric 111, volatile memory 112, persistent storage 113 (including an operating system 122 and the volumetric object adaptation engine 200, as identified above), a peripheral device set 114 (including a user interface (UI), a device set 123, storage 124, and an Internet of Things (IoT) sensor set 125), and a network module 115. The remote server 104 includes a remote database 130. The public cloud 105 includes a gateway 140, a cloud orchestration module 141, a set of host physical machines 142, a set of virtual machines 143, and a set of containers 144.

[0018] Computer 101 may take the form of a desktop computer, a laptop computer, a tablet computer, a smartphone, a smartwatch or other wearable computer, a mainframe computer, a quantum computer, or any other form of computer or mobile device now known or later developed that is capable of executing programs, accessing a network, or querying a database, such as remote database 130. As is well understood in the field of computer technology, and depending on the technology, execution of a computer-implemented method may be distributed among multiple computers and / or multiple locations. However, in this presentation of computing environment 100, to keep the presentation as concise as possible, the detailed discussion focuses on a single computer, specifically computer 101. Although computer 101 is not depicted in FIG. 1 within a cloud, it may be located within a cloud. However, computer 101 is not required to reside within a cloud except to any extent that may be expressly indicated.

[0019] Processor set 110 includes one or more computer processors of any type now known or to be developed in the future. Processing circuitry 120 may be distributed across multiple packages, e.g., multiple tailored integrated circuit chips. Processing circuitry 120 may implement multiple processor threads and / or multiple processor cores. Cache 121 is memory located within the processor chip package and is typically used for data or code that should be available for fast access by threads or cores executing on processor set 110. Cache memory is typically organized into multiple levels depending on relative proximity to the processing circuitry. Alternatively, some or all caches for a processor set may be located “off-chip.” In some computing environments, processor set 110 may be designed to operate with qubits and perform quantum computing.

[0020] Computer-readable program instructions are typically loaded onto computer 101 and cause processor set 110 of computer 101 to execute a series of operational steps, thereby enabling a computer-implemented method, such that the instructions so executed instantiate the methods specified in the computer-implemented method flowcharts and / or descriptions contained herein (collectively referred to as the "methods of the present invention"). These computer-readable program instructions are stored in various types of computer-readable storage media, such as cache 121 and other storage media discussed below. The program instructions and associated data are accessed by processor set 110 to control and direct the execution of the methods of the present invention. In computing environment 100, at least a portion of the instructions for executing the methods of the present invention may be stored in dynamic system training data generation engine 200 in persistent storage 113.

[0021] Communications fabric 111 is the signal-conducting pathway that allows various components of computer 101 to communicate with one another. Typically, this fabric is made up of switches and conductive pathways, such as switches and conductive pathways that make up buses, bridges, physical input / output ports, and the like. Other types of signal communication pathways may be used, such as fiber optic communication pathways and / or wireless communication pathways.

[0022] Volatile memory 112 may be any type of volatile memory now known or later developed. Examples include dynamic random access memory (RAM) or static RAM. Typically, volatile memory 112 is characterized by random access, although this is not required unless expressly indicated. In computer 101, volatile memory 112 is located in a single package and is internal to computer 101; however, alternatively or additionally, volatile memory may be distributed across multiple packages and / or located external to computer 101.

[0023] The persistent storage 113 may be any form of non-volatile storage for a computer, now known or later developed. The non-volatility of this storage means that stored data is maintained regardless of whether power is supplied to the computer 101 and / or to the persistent storage 113 directly. While the persistent storage 113 may be read-only memory (ROM), typically at least a portion of the persistent storage allows data to be written, data to be deleted, and data to be rewritten. Some well-known forms of persistent storage include magnetic disks and solid-state storage devices. The operating system 122 may take several forms, including various known proprietary operating systems or open-source Portable Operating System Interface-type operating systems employing a kernel. The code included in the volumetric object adaptation engine 200 typically includes at least a portion of the computer code involved in performing the methods of the present invention.

[0024] The peripheral device set 114 includes a set of peripheral devices of the computer 101. Data communication connections between the peripheral devices and other components of the computer 101 may be implemented in various ways, such as Bluetooth® connections, Near-Field Communication (NFC) connections, connections made by cable (such as a Universal Serial Bus (USB)-type cable), insertion-type connections (e.g., a Secure Digital (SD) card), connections made through a local area communication network, and even connections made through a wide area network such as the Internet. In various embodiments, the UI device set 123 may include components such as a display screen, speakers, microphones, wearable devices (such as goggles and smartwatches), keyboards, mice, printers, touchpads, game controllers, and haptic devices. The storage 124 may be external storage such as an external hard drive or insertable storage such as an SD card. The storage 124 may be persistent and / or volatile. In some embodiments, the storage 124 may take the form of a quantum computing storage device for storing data in the form of qubits. In embodiments where computer 101 is required to have a large amount of storage (e.g., where computer 101 stores and manages a large database locally), this storage may be provided by a peripheral storage device designed to store very large amounts of data, such as a storage area network (SAN) shared by multiple, geographically distributed computers. IoT sensor set 125 is made up of sensors that can be used in Internet of Things applications. For example, one sensor may be a thermometer and another sensor may be a motion detector.

[0025] Network module 115 is a collection of computer software, hardware, and firmware that enables computer 101 to communicate with other computers over WAN 102. Network module 115 may include hardware such as a modem or Wi-Fi® signal transceiver, software for packetizing and / or depacketizing data for communication network transmission, and / or web browser software for communicating data over the Internet. In some embodiments, the network control and network forwarding functions of network module 115 are performed on the same physical hardware device. In other embodiments (e.g., embodiments utilizing Software-Defined Networking (SDN)), the control and forwarding functions of network module 115 are performed on physically separate devices, such that the control function manages multiple different network hardware devices. Computer-readable program instructions for implementing the methods of the present invention may be downloaded to computer 101 from an external computer or external storage device, typically through a network adapter card or network interface included in network module 115.

[0026] WAN 102 is any wide area network (e.g., the Internet) capable of communicating computer data over non-local distances using any technology for communicating computer data now known or later developed. In some embodiments, WAN 102 may be replaced and / or supplemented by a local area network (LAN) designed to communicate data between devices located in a local area, such as a Wi-Fi network. WANs and / or LANs typically include copper transmission cables, optical fiber transmissions, wireless transmissions, and computer hardware such as routers, firewalls, switches, gateway computers, and edge servers.

[0027] End-user device (EUD) 103 is any computer system used and controlled by an end user (e.g., a customer of the enterprise that operates computer 101) and may take any of the forms described above in connection with computer 101. EUD 103 typically receives useful and useful data from the operation of computer 101. For example, in the hypothetical case where computer 101 is designed to provide recommendations to the end user, the recommendations would typically be communicated from network module 115 of computer 101 over WAN 102 to EUD 103. In this manner, EUD 103 can display or otherwise present the recommendations to the end user. In some embodiments, EUD 103 may be a client device, such as a thin client, a heavy client, a mainframe computer, a desktop computer, and the like.

[0028] Remote server 104 is any computer system that provides at least some data and / or functionality to computer 101. Remote server 104 may be controlled and used by the same entity that operates computer 101. Remote server 104 represents a machine that collects and stores useful and useful data for use by other computers, such as computer 101. For example, in the hypothetical case where computer 101 is designed and programmed to provide recommendations based on past data, then this past data may be provided to computer 101 from remote database 130 of remote server 104.

[0029] A public cloud 105 is any computer system available for use by multiple entities that provides on-demand availability of computer system resources and / or other computer functionality, particularly data storage (cloud storage) and computing power, without direct active management by users. Cloud computing typically leverages resource sharing to achieve coherence and economies of scale. Direct active management of public cloud 105 computing resources is performed by computer hardware and / or software in cloud orchestration module 141. The computing resources provided by public cloud 105 are typically implemented by virtual computing environments running on various computers comprising host physical machine set 142, which is the universe of physical computers within and / or available in public cloud 105. Virtual computing environments (VCEs) typically take the form of virtual machines from virtual machine set 143 and / or containers from container set 144. It is understood that these VCEs can be stored as images and can be transferred among and between various physical machine hosts, either as images or after instantiation of the VCEs. Cloud orchestration module 141 manages the transfer and storage of images, deploys new instantiations of VCE, and manages active instantiations of VCE deployments. Gateway 140 is a collection of computer software, hardware, and firmware that enables public cloud 105 to communicate over WAN 102.

[0030] Some further discussion of virtualized computing environments (VCEs) is now provided. A VCE can be stored as an "image." A new, active instance of a VCE can be instantiated from an image. Two well-known types of VCEs are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to a feature of an operating system in which the kernel allows the existence of multiple isolated user space instances, called containers. These isolated user space instances typically behave as actual computers from the perspective of programs running within them. A computer program running on a typical operating system can utilize all of the computer's resources, such as connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, a program running inside a container can only use the contents of the container and of the devices assigned to the container; this feature is known as containerization.

[0031] A private cloud 106 is similar to a public cloud 105, except that its computing resources are available only for use by a single enterprise. While the private cloud 106 is shown in communication with the WAN 102, in other embodiments, the private cloud may be completely disconnected from the Internet and accessible only through a local / private network. A hybrid cloud is a composite of multiple clouds of different types (e.g., private, community, or public cloud types), often each implemented by a different vendor. While each of the multiple clouds remains a separate, discrete entity, the larger hybrid cloud architecture is bound together by standardized or proprietary technologies that enable orchestration, management, and / or data / application portability between the constituent clouds. In this embodiment, both the public cloud 105 and the private cloud 106 are part of a larger hybrid cloud.

[0032] Figure 2A is a block diagram of a computer system 210 according to one embodiment of the present invention. As shown in Figure 2, the computer system 210 includes a server 212. Running on the server 212 is a dynamic system training data generation engine 200 (described in more detail in Figure 2B). The server 212 is connected to a network 218 and a dynamic function dictionary 216. Connected to the network 218 is a virtual reality system 220.

[0033] The dynamic function dictionary 216 includes a database of functions. The functions may be categorized based on dimensionality (e.g., 2-D, 3-D, 4-D, etc.). The functions may be based on past or previous dynamical systems. In one embodiment, the dynamical system may be a known model system. For example, the system may be a McKee-Glass time series, a Rossler system, a van der Pol oscillator, a Lotka-Volterra, a Henon-Heiles, and a Lorentz, to name a few. The system may be stored as a mathematical representation, such as a vector representation. In one embodiment, a dynamical system encoding model may be an entity that encodes the functions in the dynamic function dictionary 216.

[0034] FIG. 2B shows a block diagram of the dynamic system training data generation engine 200. Shown operating on the dynamic system training data generation engine 200 are a dynamic system encoding module 232, a contrastive learning module 234, and a dynamic system decoding module 236. In one embodiment, the dynamic system training data generation engine 200 can generate dynamic system training data from sample data of a dynamic system. Additionally, the dynamic system training data generation engine 200 can identify, from the dynamic function dictionary 216, a dynamic system similar to that of the sample data of the dynamic system. Furthermore, the dynamic system training data generation engine 200 can encode the sample data from the dynamic system into a mathematical representation. For example, sample physical data from the dynamic system can be provided to the dynamic system training data generation engine 200. The sample physical data can be encoded and analyzed by the dynamic system training data generation engine 200. Functions having similar mathematical expressions can be identified based on contrastive learning. Using the identified function, training data can be generated from a diffusion decoder, where the diffusion decoder uses random noise along with the mathematical representation of the sample data to generate training data for downstream purposes, such as future prediction.

[0035] The dynamic system encoding module 232 is a computer module that can receive sample dynamic system data and convert the sample into a mathematical representation. The sample data can be time-series data containing all measurements and readings of a dynamic system. The dynamic system can be an industrial production system associated with chemical production or natural gas or oil production. The dynamic system can also be associated with power production or power grid monitoring. The dynamic system encoding module can be a model composed of bi-directional recurrent neural networks (Bi-RNNs). In one embodiment, the dynamic system encoding module 232 has a particular type of RNN called a gated recurrent unit ("GRU") that controls the flow of information between different time steps.

[0036] In another embodiment, the RNN may exist in a single-layer or multi-layer architecture, with the RNNs interconnected within each layer to provide information about each time-series data point within a sample. A multi-layer RNN is one in which multiple RNN layers are stacked on top of each other. This can improve performance and help capture more complex patterns within sequential data.

[0037] The contrastive learning module 234 is a computer module that can be used to identify similar functions in a latent space. In one embodiment, the constraint learning module 234 utilizes a method used in machine learning, particularly for unsupervised learning tasks, that involves teaching a model to learn meaningful representations of data by closely comparing similar examples and pushing apart dissimilar examples. Furthermore, contrastive learning can be utilized by building a dictionary of dynamic systems by comparing them with a dynamic system function dictionary. Contrastive representation learning is a method used in machine learning, particularly for unsupervised learning tasks, that involves teaching a model to learn meaningful representations of data by closely comparing similar examples and pushing apart dissimilar examples.

[0038] This technique is often used when labeled data is scarce or expensive to acquire. The idea is to train a model on a large dataset where each example consists of two views, such as images of the same object taken at different angles, or a time series from the same system. By forcing similar examples to be close together and dissimilar examples to be far apart in the learned representation space, the model learns useful features that can later be used for tasks such as classification or prediction.

[0039] In one embodiment, the contrastive learning module 234 can distinguish multiple dynamical systems within a latent space that provides rich and diverse representations of dynamical systems (e.g., time series data), so that when provided with a sample dynamical system dataset, it can be quickly and accurately classified based on its similarity to examples in the trained dictionary. Furthermore, the contrastive learning module 234 enables efficient data compression as well as the ability to generalize to hidden or latent information in dynamical systems. Additionally, this method potentially enables a better understanding of the underlying dynamics of complex systems by analyzing the relationships between different dynamical systems within the trained representation space. To achieve this, the model is trained using a contrastive loss function that encourages similar dynamical systems to have similar representations while pushing dissimilar representations apart. During training, pairs of dynamical systems are sampled and compared based on their similarity. If two time series are deemed similar, they will be mapped close together in the trained representation (i.e., latent) space. If they are dissimilar, the model will learn to map them further apart.

[0040] In one embodiment, the contrastive learning module 234 may use a contrastive loss as a loss function. In contrastive loss, a dataset may have three data points: an anchor, a positive, and a negative. The anchor is a central data point on which the model is desired to focus. There is a positive, which is a data point that is similar to the anchor. Finally, there is a negative data point that is different from the anchor. The goal is to bring the anchor and the positive data points closer to each other in the embedding space, while moving the anchor and the negative data points further away from each other. In one embodiment, contrastive loss typically uses a distance metric (such as Euclidean distance) to measure the similarity between embeddings. The function penalizes situations where the distance between the anchor and the positive is smaller than the distance between the anchor and the negative. Essentially, contrastive loss encourages the model to learn representations where similar things are close together and dissimilar things are far apart. This allows the contrastive learning module 234 to focus on relationships by contrasting examples rather than simply labeling them. Furthermore, it is suitable for large datasets: it performs well on large unlabeled datasets where accurate labels are difficult to obtain.

[0041] The trained model can then be used to classify new sample dynamic system datasets by finding their nearest neighbors in the learned representation space and making predictions based on their labels. This approach allows for efficient data compression and the ability to generalize to unknown systems, making it useful in a variety of applications such as prediction, control, and system identification in fields such as physics, engineering, finance, and biology.

[0042] The dynamic system decoding module 236 is a computer program capable of decoding the trained representation space and generating multiple versions of the input sample dynamic system dataset. In one embodiment, the dynamic system decoding module 236 receives a representation of the input dataset and decodes the encoded sample dataset to generate multiple versions of the input sample dynamic system dataset. For example, the dynamic system decoding module 236 can be trained via a denoising stochastic diffusion model. The training of the denoising stochastic diffusion model consists of a forward diffusion process and a reverse diffusion process. In the forward diffusion process, a clean sample dataset is gradually corrupted by adding Gaussian noise for multiple steps until the dataset becomes pure noise. Next, in the reverse diffusion process, the denoising stochastic diffusion model is trained to extract the original data from the pure noise. The dynamic system decoding module 236 learns to predict the added noise at each step, essentially learning how to reverse the diffusion process. This allows the dynamic system decoding module 236 to generate dynamic system data from random noise. For example, given a representation of sample data, the dynamic system decoding module 236 first merges the representation with random noise and then iteratively reduces the noise through a de-diffusion process until a clean data set is reconstructed. This allows for a generative function that starts with pure noise and iteratively de-noises it, allowing the diffusion model to generate new, unknown data samples.

[0043] Reference is now made to Figure 3, which is a flowchart illustrating steps of the present invention, generally designated 300. In step 302, the contrastive learning module 234 may classify multiple dynamic systems based on the contrastive learning. For example, the contrastive learning may determine the embedding of functions within a dynamic function dictionary.

[0044] In step 304, the diffusion decoder is trained. For example, the dynamic system decoding module may be trained using a function from the dynamic function dictionary 216. The function may gradually have noise added to it, and the noise may be removed by the dynamic system decoding module 216 over multiple iterations until a satisfactory result is returned (i.e., the dynamic function appears to be what it was before the noise was added to it). Furthermore, the dynamic system decoding module 236 may be required to receive the dynamic function as an encoding and decode the noisy encoding back to the original function.

[0045] In step 306, a first time series data segment may be provided to the dynamic system encoding module 232. For example, a time series data set for a dynamic system such as an industrial chemical paint mixing system. The time series may consist of multiple data points in different modes and measurements, such as temperature, pressure, weight, spectroscopy readings, etc.

[0046] At step 308, the dynamic system training data generation engine 200 may generate multiple dynamic system training data based on the provided time series data set. For example, the provided data set may be encoded by the dynamic system encoding module 232. The encoding may be compared to the dynamic functions in the dynamic function dictionary 216 by the contrastive learning module 234. Once similar dynamic functions are identified, the dynamic system decoding module 216 may generate multiple data points from the encoding using random noise.

[0047] FIG. 4 is a block diagram of information flow through an exemplary embodiment of the present invention, generally designated 400. 402 illustrates an exemplary input time series of a dynamic system. 404 illustrates an exemplary embodiment of the architecture of the dynamic system encoding module 232. The reader will note the single layer of interconnected GRUs. This is for illustrative purposes, as there may be more layers of GRUs and / or more GRUs within a layer. 406 illustrates a simplified embedding using a dynamic function from the dynamic function dictionary. At 406, the contrastive learning module 234 determines or compares the embedding of the input time series with that of the embedding from the dynamic function dictionary to classify the input time series. 408 illustrates an exemplary architecture of the dynamic system decoding module 236. At 408, the dynamic system decoding module 236 receives the embedding and classification and may generate new sample time series 410 to train a time series prediction model (not shown).

[0048] While the descriptions of various embodiments of the present invention have been presented for illustrative purposes, they are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein has been selected to best explain the principles, practical applications, or technical improvements of the embodiments over technologies found in the marketplace, or to enable others skilled in the art to understand the embodiments disclosed herein.

[0049] The computer-implemented method includes providing a plurality of dynamic systems to a dynamic system dictionary, where the dynamic system dictionary has a library of functions; classifying each of the plurality of dynamic systems, where the classifying includes generating hierarchical dynamic system data based on constraint learning including an encoder and a noise generator; training a diffusion decoder to generate time series segments based on the classified plurality of dynamic systems; providing a first time series data segment; and generating time series dynamic training data based on the diffusion decoder using the first time series data segment as an input.

[0050] Embodiments of the computer-implemented method may also include the dynamic system dictionary being based on a multi-dimensional classification of associated systems.

[0051] The computer-implemented method embodiment may further include an aspect where classifying each of the plurality of dynamical systems further comprises receiving an input time series, encoding the input time series based on a plurality of recurrent neural networks, where the input time series is encoded into a multidimensional vector, determining a first similar dynamical system for the encoded input time series from the dynamical systems dictionary based on contrastive learning, and decoding the input time series based on a plurality of recurrent neural network decoding units.

[0052] Embodiments of the computer-implemented method may further include the aspect where the diffusion decoder generates time-series dynamic training data using a representation of sample data and random noise.

[0053] Embodiments of the computer-implemented method may further include an aspect in which the dynamical systems dictionary comprises a set of multiple dynamical systems having various dimensions, such as a set including Rossler systems, McKee-Glass time series, van der Pol systems, Lorenz systems, Lotka-Volterra systems, and Henon-Heiles systems.

[0054] Embodiments of the computer-implemented method may further include an aspect where the first time-series data segment is a physical system.

[0055] The computer-implemented method embodiment may further comprise an aspect where the encoding step is further based on a multi-layer bidirectional gated recurrent unit.

[0056] Further, embodiments of the present invention may include a computer system for generating time-series dynamic system training data, the system including a memory and a processor in communication with the memory. The processor may be configured to perform operations. The operations may include providing a plurality of dynamic systems to a dynamic system dictionary, where the dynamic system dictionary includes a library of functions. The operations may also include classifying each of the plurality of dynamic systems, where the classifying step includes generating hierarchical dynamic system data based on constraint learning including an encoder and a noise generator. The operations may also include training a diffusion decoder to generate time-series segments, where the decoder is based on the classified plurality of dynamic systems. The operations may include providing a first time-series data segment and generating time-series dynamic training data based on the diffusion decoder using the first time-series data segment as an input.

[0057] Embodiments of the computer system may include an aspect in which the dynamic system dictionary is based on a multi-dimensional classification of associated systems.

[0058] An embodiment of the computer system may include an aspect where the procedure for classifying each of the plurality of dynamical systems further includes operations of receiving an input time series; encoding the input time series based on a plurality of recurrent neural networks, where the input time series is encoded into a multi-dimensional vector; determining a first similar dynamical system for the encoded input time series from the dynamical system dictionary based on contrastive learning; and decoding the input time series based on a plurality of recurrent neural network decoding units.

[0059] An embodiment of the computer system may include an aspect in which the diffusion decoder generates time-series dynamic training data using a representation of sample data and random noise.

[0060] Embodiments of the computer system may include an aspect in which the dynamical systems dictionary is comprised of a set of multiple dynamical systems having various dimensions, such as a set including Rossler systems, McKee-Glass time series, van der Pol systems, Lorenz systems, Lotka-Volterra systems, and Henon-Heiles systems.

[0061] Embodiments of the computer system may include an aspect where the first time-series data segment is a physical system.

[0062] Embodiments of the computer system may include an aspect where the encoding procedure is further based on a multi-layer bidirectional gated recurrent unit.

[0063] Further, embodiments of the present invention may include a computer program product for generating time-series dynamic system training data, the computer program product comprising: a computer storage device; and program instructions stored on the computer storage device, the program instructions including: program instructions for providing a plurality of dynamic systems to a dynamic system dictionary, the dynamic system dictionary including a library of functions; program instructions for classifying each of the plurality of dynamic systems, the classifying step including generating hierarchical dynamic system data based on constraint learning including an encoder and a noise generator; program instructions for training a diffusion decoder to generate time-series segments based on the classified plurality of dynamic systems; program instructions for providing a first time-series data segment; and program instructions for generating time-series dynamic training data based on the diffusion decoder using the first time-series data segment as an input.

[0064] Embodiments of the computer program product may comprise an aspect where the dynamic system dictionary is based on a multi-dimensional classification of associated systems.

[0065] An embodiment of the computer system may include an aspect where the procedure for classifying each of the plurality of dynamical systems further includes program instructions for receiving an input time series; program instructions for encoding the input time series based on a plurality of recurrent neural networks, where the input time series is encoded into a multi-dimensional vector; program instructions for determining a first similar dynamical system for the encoded input time series from the dynamical systems dictionary based on contrastive learning; and program instructions for decoding the input time series based on a plurality of recurrent neural network decoding units.

[0066] An embodiment of the computer program product may comprise the diffusion decoder generating time series dynamic training data using a representation of sample data and random noise.

[0067] Embodiments of the computer program product may include an aspect in which the dynamical systems dictionary is comprised of a set of multiple dynamical systems having various dimensions, for example, a set including Rossler systems, McKee-Glass time series, van der Pol systems, Lorenz systems, Lotka-Volterra systems, and Henon-Heiles systems.

[0068] Embodiments of the computer program product may include an aspect where the first time-series data segment is a physical system.

Claims

1. 1. A computer-implemented method for generating time-series dynamic system training data, comprising: providing a plurality of dynamic systems in a dynamic systems dictionary, wherein said dynamic systems dictionary includes a library of functions; classifying each of the plurality of dynamic systems, wherein the classifying step comprises generating hierarchical dynamic system data based on constraint learning, the hierarchical dynamic system data including an encoder and a noise generator; training a diffusion decoder to generate time series segments based on the classified plurality of dynamical systems; providing a first time-series data segment; and generating time-series dynamic training data based on the diffusion decoder using the first time-series data segment as an input; A computer-implemented method comprising:

2. The computer-implemented method of claim 1 , wherein the dynamic system dictionary is based on a multi-dimensional classification of associated systems.

3. The step of classifying each of the plurality of dynamic systems comprises: receiving an input time series; encoding the input time series based on a plurality of recurrent neural networks, wherein the input time series is encoded into a multi-dimensional vector; determining a first similar dynamic system for the encoded input time series from the dynamic system dictionary based on contrastive learning; and decoding the input time series based on a plurality of recurrent neural network decoding units; The computer-implemented method of claim 1 or 2, further comprising:

4. The computer-implemented method of claim 1 or 2, wherein the diffusion decoder generates time-series dynamic training data using a representation of sample data and random noise.

5. 3. The computer-implemented method of claim 1 or 2, wherein the dynamical systems dictionary comprises a set of a number of dynamical systems with various dimensions, such as a set including Rossler systems, McKee-Glass time series, van der Pol systems, Lorenz systems, Lotka-Volterra systems, and Henon-Heiles systems.

6. The computer-implemented method of claim 1 or 2, wherein the first time-series data segment is a physical system.

7. The computer-implemented method of claim 3 , wherein the encoding step is further based on multi-layer bidirectional gated recurrent units.

8. 1. A computer system for generating time series dynamic system training data, comprising: memory; and a processor in communication with said memory wherein the processor performs operations to: providing a plurality of dynamic systems in a dynamic systems dictionary, wherein said dynamic systems dictionary includes a library of functions; classifying each of the plurality of dynamic systems, wherein the classifying step includes a step of generating hierarchical dynamic system data based on constraint learning, the step including an encoder and a noise generator; training a diffusion decoder to generate time series segments based on the classified plurality of dynamical systems; providing a first time series data segment; and generating time-series dynamic training data based on the diffusion decoder using the first time-series data segment as an input; 1. A computer system configured to:

9. The computer system of claim 8 , wherein the dynamic system dictionary is based on a multi-dimensional classification of associated systems.

10. The procedure for classifying each of the plurality of dynamic systems comprises: Receives an input time series; encoding the input time series based on a plurality of recurrent neural networks, where the input time series is encoded into a multi-dimensional vector; determining a first similar dynamic system for the encoded input time series from the dynamic system dictionary based on contrastive learning; and Decoding the input time series based on a plurality of recurrent neural network decoding units.

10. The computer system of claim 8 or 9, further comprising:

11. 10. The computer system of claim 8, wherein the diffusion decoder generates time-series dynamic training data using a representation of sample data and random noise.

12. 10. The computer system of claim 8 or 9, wherein the dynamical systems dictionary comprises a set of a number of dynamical systems with various dimensions, such as a set including Rossler systems, McKee-Glass time series, van der Pol systems, Lorenz systems, Lotka-Volterra systems, and Henon-Heiles systems.

13. 10. The computer system of claim 8 or 9, wherein the first time-series data segment is a physical system.

14. 11. The computer system of claim 10, wherein the encoding procedure is further based on multi-layer bidirectional gated recurrent units.

15. 1. A computer program for generating time series dynamic system training data, the computer program comprising: providing a plurality of dynamic systems to a dynamic systems dictionary, said dynamic systems dictionary including a library of functions; a step of classifying each of the plurality of dynamic systems, wherein the step of classifying includes a step of generating hierarchical dynamic system data based on constraint learning, the step including an encoder and a noise generator; training a diffusion decoder to generate time series segments based on the classified plurality of dynamical systems; providing a first time-series data segment; and generating time-series dynamic training data based on the diffusion decoder using the first time-series data segment as input; A computer program for executing

16. The computer program of claim 15 , wherein the dynamic system dictionary is based on a multi-dimensional classification of associated systems.

17. The procedure for classifying each of the plurality of dynamic systems comprises: receiving an input time series; encoding the input time series based on a plurality of recurrent neural networks, wherein the input time series is encoded into a multi-dimensional vector; determining a first similar dynamic system for the encoded input time series from the dynamic system dictionary based on contrastive learning; and decoding the input time series based on a plurality of recurrent neural network decoding units; 17. The computer program of claim 15 or 16, further comprising:

18. 17. The computer program of claim 15 or 16, wherein the diffusion decoder generates time-series dynamic training data using a representation of sample data and random noise.

19. 17. The computer program of claim 15 or 16, wherein the dynamical systems dictionary comprises a set of a number of dynamical systems with various dimensions, such as a set including Rossler systems, McKee-Glass time series, van der Pol systems, Lorenz systems, Lotka-Volterra systems, and Henon-Heiles systems.

20. 17. The computer program of claim 15 or 16, wherein the first time-series data segment is a physical system.