Predictively Robust Model Training

By embedding and perturbing time series datasets to train a learning function, the method addresses data drift issues, enhancing model robustness and reducing retraining frequency.

JP2025520976AActive Publication Date: 2025-07-03NEC CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025500763
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-07-12
Filing Date
2023-04-21
Publication Date
2025-07-03
Estimated Expiration
2043-04-21

AI Technical Summary

Technical Problem

Existing supervised machine learning models face challenges in maintaining accuracy and effectiveness due to data drift and environmental changes, requiring frequent retraining, which is costly and inefficient.

Method used

A method involving embedding time series datasets into feature vectors, predicting future feature vectors, creating perturbed future datasets, and training a learning function to generate a predictively robust model that accounts for data drift and uncertainty.

Benefits of technology

The method extends the lifespan of the model, reducing the need for frequent retraining and associated costs by ensuring the model remains effective in varying data conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025520976000001_ABST
    Figure 2025520976000001_ABST
Patent Text Reader

Abstract

A predictively robust model is trained by performing: embedding the distribution of each time series dataset among a plurality of time series datasets into a feature vector; predicting a future feature vector of the distribution of a future dataset based on the feature vectors of each time series dataset among the plurality of time series datasets; creating a future dataset from the future feature vector; perturbing the future dataset to generate a plurality of perturbed future datasets; and training a learning function using the future dataset and each perturbed future dataset to generate the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a computer-readable medium, a method, and an apparatus for training a predictively robust model.

Background Art

[0002] In supervised machine learning, training is based on a training dataset curated by those proficient in the process. The curation of the training dataset can be a large-scale and costly process involving a lot of man-hours. Once the model is trained by the training dataset, more man-hours may be spent to verify the trained model before implementation. After implementation, the performance of the trained model is monitored for accuracy and effectiveness. The model is retrained when the accuracy or effectiveness is no longer appropriate. Even when the model is carefully trained and verified, the accuracy or effectiveness will ultimately lose its validity due to data drift, environmental changes, etc. When using the model in some applications, the issue is not whether the model is retrained, but when it is retrained.

Summary of the Invention

[0003] According to an exemplary aspect of the present disclosure, a computer-readable medium includes causing a computer to perform operations including embedding the distribution of each time-series dataset among a plurality of time-series datasets into a feature vector, predicting a future feature vector of the distribution of a future dataset based on the feature vectors of each time-series dataset among the plurality of time-series datasets, creating a future dataset from the future feature vector, perturbing the future dataset to generate a plurality of perturbed future datasets, and training a learning function using the future dataset and each perturbed future dataset to generate a model.

[0004] According to an exemplary aspect of the present disclosure, the method includes embedding the distribution of each time series dataset among a plurality of time series datasets into a feature vector, predicting a future feature vector of the distribution of a future dataset based on the feature vectors of each time series dataset among the plurality of time series datasets, creating a future dataset from the future feature vector, perturbing the future dataset to generate a plurality of perturbed future datasets, and training a learning function using the future dataset and each perturbed future dataset to generate a model.

[0005] According to an exemplary aspect of the present disclosure, the apparatus includes a controller including a circuit configured to perform embedding the distribution of each time series dataset among a plurality of time series datasets into a feature vector, predicting a future feature vector of the distribution of a future dataset based on the feature vectors of each time series dataset among the plurality of time series datasets, creating a future dataset from the future feature vector, perturbing the future dataset to generate a plurality of perturbed future datasets, and training a learning function using the future dataset and each perturbed future dataset to generate a model.

Brief Description of the Drawings

[0006] Aspects of the present disclosure are preferably understood from the following detailed description when read in conjunction with the accompanying drawings. Note that various features are not drawn to scale in accordance with standard industry practice. In fact, the dimensions of various features may be arbitrarily increased or decreased for clarity of explanation.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

DETAILED DESCRIPTION OF THE INVENTION

[0007] The following disclosure provides many different embodiments or examples for implementing various features of the provided subject matter. In the following, specific examples of components, values, operations, materials, arrangements, etc. are described in order to simplify the present disclosure. Of course, these are merely examples and are not intended to be limiting. Other components, values, operations, materials, arrangements, etc. are conceivable. Further, the present disclosure may repeat reference numbers and / or letters in various embodiments. This repetition is for simplicity and clarity and does not in itself define the relationship between the various embodiments and / or configurations described.

[0008] In data classification, algorithms are used to divide a dataset into multiple classes. These classes may have multiple sub-populations or sub-categories that are not related to the immediate classification task. Some sub-populations or sub-categories are frequent and some are infrequent. The relative frequency of the sub-populations can affect the performance of a classifier, which is an algorithm used to classify the data of a dataset into multiple classes. Some classifiers are trained using a concept known as the following Empirical Risk Minimization (ERM), [Number] where $\hat{h}$ is the trained classifier algorithm, $l$ is the loss function, $h$ θ is the classifier learning function, $x$ i is the input to the classifier function, and $h$ θ $(x$ i ) represents the class output from the classifier function, and $y$ i is the true class. However, ERM is optimized for the training dataset and does not consider the uncertainty of the training dataset nor data drift. As a result, if there is a shift in the relative frequency of the sub-populations, the performance of the classifier will degrade.

[0009] Some classification algorithms complement the training dataset with several synthetic datasets generated by perturbing the training dataset that represents the current state of the data, such as by using the following adversarial weighting scheme.

Number

Number

Number

[0010] Some algorithms consider historical data, estimate the data drift trend, and predict future datasets.

[0011] In at least some embodiments described herein, classifiers and other models are generated through predictively robust model training, taking into account data drift and uncertainty in the training dataset. In at least some embodiments, time series data is used to predict future states, which, when supplemented with perturbations of the distribution or density function of the future state and used to train the model, creates a training dataset that results in a predictively robust model. In at least some embodiments, the resulting predictively robust model exhibits a longer lifespan than a model trained using a classification algorithm that perturbs the training dataset representing the current state of the data, because the actual future state is likely to fall within a range of divergence, sometimes called a "divergence ball," centered around the predicted state rather than the current state. Since the actual future state is likely to fall within the range of divergence centered around the predicted state, at least some embodiments use a divergence smaller than the divergence centered around the current state, which reduces the likelihood of unrealistic subpopulation frequencies and further extends the lifespan of the model.

[0012] In at least some embodiments, the classifier is trained to perform well on subpopulations that are of low frequency during training. In at least some embodiments, predictively robust model training improves the lifespan of the model, which reduces the cost of model retraining, such as the labor involved in reducing the number of models in the archive, compliance, quality control, training dataset curation, and the computational resources required for model retraining.

[0013] FIG. 1 is an operational flow for predictively robust model training according to at least some embodiments of the present disclosure. The operational flow provides a method for predictively robust model training. In at least some embodiments, one or more operations of the method are performed by a controller of a device that includes sections for performing specific operations, such as the controller and device shown in FIG. 13 described below.

[0014] In S100, the controller or its section groups time series data into data sets. In at least some embodiments, the controller groups the time series data into a plurality of time series data sets. In at least some embodiments, the time series are grouped into equally spaced time steps. In at least some embodiments, each group represents the historical training data of the model. In at least some embodiments, each group includes a distribution of data samples representing the state at the corresponding time. In at least some embodiments, the group including the distribution of the latest data sample represents the current state. In at least some embodiments, each group includes a density function representing the state at the corresponding time. In at least some embodiments, the controller receives the time series that has already been grouped and proceeds directly to the distribution data set embedding in S110.

[0015] In S110, the embedding section embeds the distribution of each data set. In at least some embodiments, the embedding section embeds the distribution of each time series data set among a plurality of time series data sets into a feature vector. In at least some embodiments, the embedding section estimates the probability density function of each time series data set. In at least some embodiments, the embedding section performs the data set distribution embedding process described later with respect to FIG. 2.

[0016] In S120, the prediction section predicts future feature vectors. In at least some embodiments, the prediction section predicts the future feature vectors of the distribution of future data sets based on the feature vectors of each time series data set among a plurality of time series data sets. In at least some embodiments, the prediction section determines the data drift trend. In at least some embodiments, the prediction section predicts future feature vectors by estimating the data drift trend indicated by the historical data. In at least some embodiments, the prediction section performs the future feature vector prediction process described later with respect to FIG. 5.

[0017] In S130, the creation section creates a future dataset. In at least some embodiments, the creation section creates a future dataset from the future feature vectors predicted in S120. In at least some embodiments, the creation section decodes the future feature vectors into a future probability density function, generates weights according to the difference between the future probability density function and the probability density function of the current state, and resamples the dataset representing the current state according to the generated weights. In at least some embodiments, the creation section performs the future dataset creation process described below with respect to FIG. 7.

[0018] In S140, the perturbation section perturbs the future dataset. In at least some embodiments, the perturbation section perturbs the future dataset to generate a plurality of perturbed future datasets. In at least some embodiments, the perturbation section complements the dataset representing the future state with a perturbation of the distribution or density function of the future state to create a training dataset that, when used to train the model, results in a predictively robust model. In at least some embodiments, the perturbation section performs the future dataset perturbation process described below with respect to FIG. 8.

[0019] In S150, the training section trains a learning function. In at least some embodiments, the training section trains the learning function using a future dataset and each perturbed future dataset to generate a model. In at least some embodiments, the training section trains the learning function to classify samples within the future dataset and each perturbed future dataset. In at least some embodiments, the learning function is a linear classifier. In at least some embodiments, the learning function is a non-linear classifier. In at least some embodiments, each sample includes a label representing a ground truth classification. In at least some embodiments, the learning function is trained to output the classification represented by the label in response to its application to the sample.

[0020] FIG. 2 is a diagram of a dataset 202 having classes and sub-populations, according to at least some embodiments of the present disclosure. In at least some embodiments, the dataset 202 is a time series dataset including a plurality of samples. Each sample is characterized by x and y coordinates and is paired with a label reflecting the class to which it belongs. The classes include a first class indicated by a + in FIG. 2 and a second class indicated by a ○ in FIG. 2. FIG. 2 shows each sample as its corresponding label and plots it at a position corresponding to the x and y coordinates of the characteristics of the sample.

[0021] The first-class dataset 202 has two visible sub-populations shown as sub-population 204 and sub-population 205. Sub-population 204 has many samples, while sub-population 205 has only five samples. It should be understood that sub-population 204 and sub-population 205 are not represented by the information provided in dataset 202. Instead, sub-population 204 and sub-population 205 may have some commonality in the original data that constitutes dataset 202 or from which dataset 202 is formed, but such commonality is not actually represented in the information provided by the dataset. Thus, sub-population 205 may exist by mere chance without any commonality. On the other hand, sub-population 205 may not fully represent the actual commonality. In at least some embodiments, there is no need to confirm whether sub-population 205, or any other sub-population of dataset 202, actually has commonality.

[0022] The first-class dataset 202 has noisy samples 207. The noisy samples 207 are labeled in the first class but are surrounded only by samples from the second class. The noisy samples 207 are considered to be noisy samples not because they are mislabeled, but rather because they are not useful for the process of generating a classification model. In other words, even if a classification model is trained to correctly label sample 207, such a classification model is likely to be considered "overfitted" and thus not accurate for classifying data outside of dataset 202.

[0023] FIG. 3 is an operation flow for dataset distribution embedding according to at least some embodiments of the present disclosure. The operation flow provides a method for dataset distribution embedding. In at least some embodiments, one or more operations of this method are performed by an embedding section of a device, such as the device shown in FIG. 13 described below.

[0024] In S312, the embedding section or its sub-section estimates the density function of the dataset. In at least some embodiments, as the iteration of the operation flow progresses, the embedding section estimates the density function of each time-series dataset among a plurality of time-series datasets. In at least some embodiments, the embedding section utilizes a parametric or non-parametric density estimator. In at least some embodiments, the embedding section estimates the point density function of each time-series dataset on a weighted sum basis. In at least some embodiments, the embedding section estimates [Number] (the point density function of time-series dataset j) as a mixture of base density functions, [Number] according to the following function, [Number] denoted as, where α i represents the weight assigned to the i-th base density function, [Number] feature vector [Number] [Number] is the point density function of time-series dataset j, K is the feature vector length, x is the sample, and X is the classification. In at least some embodiments, the base density function [Number] can be calculated using a mixture model algorithm such as a Gaussian mixture model (GMM). In at least some embodiments, [Number] It can also be manually generated by a data scientist.

[0025] In S314, the embedding section or its sub-section applies an embedding function to the density function estimated in S312. In at least some embodiments, as the iteration of the operation flow progresses, the embedding section embeds the density function of each time series data set. In at least some embodiments, the embedding section is the feature vector of the density function [α1, α2,... α K into the Euclidean space. In at least some embodiments, the embedding section utilizes principal component analysis (PCA), independent component analysis (ICA), or another dimensionality reduction technique to obtain [β1, β2,... β L =[α1, α2,... α K *W so that the feature vector length is compressed from K dimensions to L dimensions [β1, β2,... β L , where K > L,

Number

[0026] In S316, the embedding section or its sub-section determines whether all data sets have been embedded. If the embedding section determines that there are remaining time series data sets that have not been embedded, the operation flow returns to the density function estimation in S312 to estimate the density function of the next time series data set (S318). If the embedding section determines that all time series data sets have been embedded into the feature vector, the operation flow ends.

[0027] In at least some embodiments, the embedding section embeds the distribution of each time series dataset without estimating a density function. In at least some embodiments, the embedding section directly embeds the distribution of each time series dataset into the feature vector.

[0028] FIG. 4 is a map 411 of feature vectors representing time series dataset distributions according to at least some embodiments of the present disclosure. Map 411 shows the feature vectors of each time series dataset, such as feature vector 415, representing the current state time series dataset mapped in a two-dimensional Euclidean space. In at least some embodiments, the embedding section embeds each time series dataset into more than two feature vectors, making visualization difficult. However, there is no need to visualize or interpret the feature vectors. Map 411 and the feature vectors mapped thereon are simplified for illustration purposes.

[0029] FIG. 5 is an operation flow of future feature vector prediction according to at least some embodiments of the present disclosure. The operation flow provides a method for predicting future feature vectors. In at least some embodiments, one or more operations of the method are performed by a prediction section of a device, such as the device shown in FIG. 13 described below.

[0030] In S522, the prediction section or a sub-section thereof initializes a trend estimator. In at least some embodiments, the trend estimator is a multivariate time series prediction learning function that learns an equation to represent future observations as a function of past observations using historical time series data. In at least some embodiments, the trend estimator is an autoregressive integrated moving average (ARIMA(p, d, q)) model. In at least some embodiments, the prediction section assigns a random value between 0 and 1 to the parameters of the trend estimator.

[0031] In S524, the prediction section or its sub-section applies a trend estimator to the feature vector. In at least some embodiments, the prediction section applies the trend estimator to the parameters [α1, α2,... α K of the feature vector. In at least some embodiments, as the iteration of the operation flow progresses, the prediction section applies the trend estimator to each feature vector.

[0032] In S525, the prediction section or its sub-section adjusts the trend estimator based on the next feature vector. In at least some embodiments, the prediction section adjusts the trend estimator by comparing the output resulting from the application to the feature vector with the parameters of the feature vector representing the subsequent time series dataset. In at least some embodiments, the feature vectors are training samples each labeled with a feature vector representing a subsequent time series dataset. In at least some embodiments, the feature vector representing the current state is not used as a training sample and is only used as a label for the feature vectors representing the preceding time series datasets.

[0033] In S526, the prediction section determines whether an end condition is satisfied. In at least some embodiments, as the iteration of the operation flow progresses, the prediction section trains the trend estimator to output a temporally subsequent feature vector in response to the application to each feature vector except the latest feature vector. In at least some embodiments, the end condition is satisfied when a predetermined number of training samples have been processed or when a predetermined number of epochs have been performed. In at least some embodiments, the end condition is satisfied when the error calculated from the loss function is less than a threshold amount. In at least some embodiments, the end condition is satisfied when the trend estimator has converged to a solution. If the end condition has not yet been satisfied, the operation flow returns to the trend estimator application in S524 and applies the next feature vector (S527). If the end condition is satisfied, the operation flow proceeds to the trained trend estimator application in S529.

[0034] In S529, the prediction section or its sub-section applies a trend estimator trained on the latest feature vector. In at least some embodiments, the prediction section applies a trend estimator to the latest feature vector to output a future feature vector. In at least some embodiments, the prediction section applies a trend estimator to the feature vector representing the current state to obtain a feature vector representing a future data set.

[0035] FIG. 6 is a map 611 showing a future feature vector 621 between time series dataset distribution feature vectors according to at least some embodiments of the present disclosure. The map 611 also shows the feature vectors of each time series dataset, such as the feature vector 615, representing the time series dataset of the current state. The map 611 is substantially the same as the map 411 of FIG. 4 in terms of structure and function unless otherwise indicated.

[0036] FIG. 7 is an operation flow for creating a future data set according to at least some embodiments of the present disclosure. The operation flow provides a method for creating a future data set. In at least some embodiments, one or more operations of the method are performed by a creation section of a device, such as the device shown in FIG. 13 described below.

[0037] In S732, the creation section or its sub-section estimates a future density function. In at least some embodiments, the creation section estimates the density function of a future data set. In at least some embodiments, the creation section applies the parameters [α1, α2,... α K of the future feature vector to EQ.5 to

Number

[0038] In S734, the creation section or its sub-section generates sample weights. In at least some embodiments, the creation section generates sample weights based on the density function of the future data set and the density function of the latest data set among a plurality of time series data sets. In at least some embodiments, the creation section generates, for each sample in the latest data set representing the current state, a sample weight w i as follows [Equation] where [Equation] is the point density function representing the latest data set, and [Equation] is the point density function representing the future data set.

[0039] In S736, the creation section or its sub-section resamples the latest data set. In at least some embodiments, the creation section resamples the latest data set according to the sample weights generated in S734. For example, w i = 3 indicates that the sample x i is three times more likely to appear in the future data set than in the current data set. Therefore, the creation section generates three samples xi in the future data set for each sample x i of the latest data set.

[0040] In at least some embodiments, the creation section directly creates the future data set from the future feature vectors.

[0041] FIG. 8 is an operation flow of future data set perturbation according to at least some embodiments of the present disclosure. The operation flow provides a method for future data set perturbation. In at least some embodiments, one or more operations of the method are performed by a perturbation section of a device, such as the device shown in FIG. 13 described below.

[0042] In S842, the perturbation section or a sub-section thereof determines the difference between the future data set and the latest data set. In at least some embodiments, the perturbation section utilizes a distance measurement algorithm to determine the distance between the future data set and the latest data set. In at least some embodiments, the perturbation section determines the difference based on the feature vectors representing the future data set and the latest data set.

[0043] In S844, the perturbation section or a sub-section thereof sets a divergence limit based on the difference between the future data set and the latest data set. In at least some embodiments, the perturbation section sets a divergence limit δ according to the difference. In at least some embodiments, the perturbation section is based on the divergence limit regarding the difference between the future data set and the latest time series data set. In at least some embodiments, the perturbation section sets the divergence limit to be greater than or equal to the difference between the future data set and the latest time series data set.

[0044] In S846, the perturbation section or a sub-section thereof generates a perturbed future data set. In at least some embodiments, the perturbation section utilizes the distributionally robust optimization (DRO) method to complement the future data set with the perturbed future data set. In at least some embodiments, the perturbation section generates a perturbed future data set by perturbing the future data set using the adversarial weighting scheme in EQ.2 and EQ.3. In at least some embodiments, each perturbed future data set diverges from the future data set within a predetermined divergence limit.

[0045] FIG. 9 is a map 911 showing the feature vectors of perturbed future data sets among the time series data set distribution feature vectors according to at least some embodiments of the present disclosure. Map 911 shows a plurality of feature vectors representing perturbed future data sets, such as feature vector 947, distributed around future feature vector 921. Map 911 also shows a boundary 945 centered on future feature vector 921 that represents the degree to which the perturbed future data set differs from the future data set. To show that the divergence limit is greater than the difference between the future data set and the latest time series data set, boundary 945 intersects feature vector 915 representing the latest data set. Map 911 is substantially the same as map 611 of FIG. 6 in structure and function unless otherwise indicated.

[0046] FIG. 10 is an operation flow for training a learning function according to at least some embodiments of the present disclosure. The operation flow provides a method for training a learning function. In at least some embodiments, one or more operations of the method are performed by a training section of an apparatus, such as the apparatus shown in FIG. 13 described below.

[0047] In S1052, the training section or a sub-section thereof initializes the learning function. In at least some embodiments, the learning function is a classification model. In at least some embodiments, the training section assigns random values between 0 and 1 to the parameters of the learning function.

[0048] In S1054, the training section or its sub-section applies a learning function to the training samples. In at least some embodiments, the training section provides the training samples as inputs to the learning function and obtains output values. In at least some embodiments, the training section provides the training samples as inputs to the learning function and obtains output classes. In at least some embodiments, the training section provides the training samples as inputs to the learning function and, for each class, obtains the probability that the training sample belongs to the class. In at least some embodiments, the training samples are selected from among samples of a future dataset and perturbed future datasets.

[0049] In S1056, the training section or its sub-section adjusts the learning function based on the labels of the training samples. In at least some embodiments, the training section compares the output values with the labels and determines the differences. In at least some embodiments, the training section applies a loss function to the output values and the labels to obtain a loss value. In at least some embodiments, the training section adjusts the weights and other parameters of the learning function based on the loss value. In at least some embodiments, the training section adjusts the weights by using gradient descent. In at least some embodiments, the training section does not adjust the learning function for each iteration of the operation flow.

[0050] In S1058, the training section determines whether an end condition is met. In at least some embodiments, as the iteration of the operation flow progresses, the training section trains a learning function to output a classification in response to the application to each training sample. In at least some embodiments, the end condition is met when a predetermined number of training samples have been processed or when a predetermined number of epochs have been performed. In at least some embodiments, the end condition is met when the loss calculated from the loss function is less than a threshold loss. In at least some embodiments, the end condition is met when the learning function has converged to a solution. If the end condition is not met, the operation flow returns to the learning function application in S1054 and applies the next training sample (S1059). If the end condition is met, the operation flow ends.

[0051] FIG. 11 is a diagram of a first classification function 1151 for a dataset 1102 having classes and subpopulations according to at least some embodiments of the present disclosure. The dataset 1102 includes a subpopulation 1104, a subpopulation 1105, and a noisy sample 1107, which respectively correspond to the subpopulation 204, the subpopulation 205, and the noisy sample 207 of FIG. 2, and thus should be understood to have the same quality unless otherwise specified.

[0052] The first classification function 1151 is plotted against the dataset 1102 and shows the decision boundary used by the first classification function 1151 to determine the classification of the samples of the dataset 1102. The first classification function 1151 has a non-linear decision boundary that is more difficult to interpret than a linear decision boundary. Whether the first classification function 1151 is easy to understand is subjective, but the non-linear decision boundary is more difficult for interested parties to understand than the linear decision boundary.

[0053] FIG. 12 is a diagram of a second classification function 1251 with respect to a dataset 1202 having classes and subpopulations, according to at least some embodiments of the present disclosure. The dataset 1202 includes a subpopulation 1204, a subpopulation 1205, and noisy samples 1207, which respectively correspond to the subpopulation 204, the subpopulation 205, and the noisy samples 207 of FIG. 2, and thus, unless otherwise specified, should be understood to have the same quality.

[0054] The second classification function 1251 is plotted against the dataset 1202 and shows the decision boundary used by the second classification function 1251 to determine the classification of samples in the dataset 1202. The second classification function 1251 has an easily understandable linear decision boundary and is thus easy to interpret, determining the classification based on which side of the decision boundary a sample falls.

[0055] FIG. 13 is a block diagram of a hardware configuration for predictive robust model training, according to at least some embodiments of the present disclosure.

[0056] An exemplary hardware configuration includes an apparatus 1360 that interacts with an input device 1369 and communicates with a network 1367. In at least some embodiments, the apparatus 1360 is integrated with the input device 1369. In at least some embodiments, the apparatus 1360 is a computer or other computing device that receives an input or command from the input device 1369. In at least some embodiments, the apparatus 1360 is a host server that connects directly to the input device 1369 or indirectly via the network 1367. In at least some embodiments, the apparatus 1360 is a computer system that includes two or more computers. In at least some embodiments, the apparatus 1360 is a computer system that executes computer-readable instructions to perform operations for physical network function device access.

[0057] Device 1360 includes a controller 1362, a memory unit 1364, a communication interface 1366, and an input / output interface 1368. In at least some embodiments, controller 1362 includes a processor or programmable circuit that executes instructions to cause a processor or programmable circuit to perform operations according to the instructions. In at least some embodiments, controller 1362 includes an analog or digital programmable circuit, or any combination thereof. In at least some embodiments, controller 1362 includes physically separated storage devices or circuits that interact via communication. In at least some embodiments, memory unit 1364 includes a non-volatile computer-readable medium that can store executable data and non-executable data for access by controller 1362 during execution of instructions. Communication interface 1366 transmits and receives data to and from network 1367. Input / output interface 1368 connects to and exchanges information with various input and output components such as input device 1369 via parallel ports, serial ports, keyboard ports, mouse ports, monitor ports, etc.

[0058] Controller 1362 includes an embedded section 1370, a prediction section 1372, a creation section 1374, a perturbation section 1376, and a training section 1378. Memory unit 1364 includes a data set 1380, a feature vector 1382, prediction parameters 1384, a future data set 1387, and a learning function 1389.

[0059] The embedding section 1370 is a circuit or instruction of the controller 1362 configured to embed the dataset distribution. In at least some embodiments, the embedding section 1370 is configured to embed the distribution of each time series dataset into a feature vector. In at least some embodiments, the embedding section 1370 utilizes information in the storage unit 1364, such as the dataset 380, and records information such as the feature vector 1382 in the storage unit 1364. In at least some embodiments, the embedding section 1370 includes subsections for implementing additional functions, as described in the aforementioned flowchart. In at least some embodiments, such subsections are referenced by names associated with the corresponding functions.

[0060] The prediction section 1372 is a circuit or instruction of the controller 1362 configured to predict future feature vectors. In at least some embodiments, the prediction section 1372 is configured to predict the future feature vectors of the distribution of future datasets based on the feature vectors of each time series dataset in the time series. In at least some embodiments, the prediction section 1372 utilizes information in the storage unit 1364, such as the feature vector 1382 and the prediction parameter 1384, and records information such as the feature vector 1382 in the storage unit 1364. In at least some embodiments, the prediction section 1372 includes subsections for implementing additional functions, as described in the aforementioned flowchart. In at least some embodiments, such subsections are referenced by names associated with the corresponding functions.

[0061] The creation section 1374 is the circuitry or instructions of the controller 1362 configured to create a future data set. In at least some embodiments, the creation section 1374 is configured to create a future data set from future feature vectors. In at least some embodiments, the creation section 1374 utilizes information from the storage unit 1364, such as the feature vector 1382, and records information, such as the future data set 1387, in the storage unit 1364. In at least some embodiments, the creation section 1374 includes subsections for implementing additional functions, as described in the flowchart above. In at least some embodiments, such subsections are referenced by names associated with the corresponding functions.

[0062] The perturbation section 1376 is the circuitry or instructions of the controller 1362 configured to perturb a data set. In at least some embodiments, the perturbation section 1376 is configured to perturb a future data set to generate a plurality of perturbed future data sets. In at least some embodiments, the perturbation section 1376 utilizes information from the storage unit 1364, such as the perturbation parameter 1386 and the future data set 1387, and records information, such as the future data set 1387, in the storage unit 1364. In at least some embodiments, the perturbation section 1376 includes subsections for performing additional functions, as described in the flowchart above. In at least some embodiments, such subsections are referenced by names associated with the corresponding functions.

[0063] The training section 1378 is a circuit or instruction of the controller 1362 configured to train a learning function. In at least some embodiments, the training section 1378 is configured to train a learning function using a future dataset and each perturbed future dataset to generate a model. In at least some embodiments, the training section 1378 utilizes information from the memory unit 1364, such as the learning function 1389. In at least some embodiments, the training section 1378 includes subsections for performing additional functions, as described in the aforementioned flowchart. In at least some embodiments, such subsections are referenced by names associated with the corresponding functions.

[0064] In at least some embodiments, the apparatus is another device capable of processing logical functions to perform the operations described herein. In at least some embodiments, the controller and the memory unit need not be completely separate devices, and in some embodiments, they share a circuit or one or more computer-readable media. In at least some embodiments, the memory unit includes a hard drive that stores both computer-executable instructions and data accessible by the controller, and the controller includes a combination of a central processing unit (CPU) and RAM, and the computer-executable instructions can be copied in whole or in part by the CPU for execution during the performance of the operations described herein.

[0065] In at least some embodiments where the apparatus is a computer, a program installed on the computer can cause the computer to function as the apparatus of the embodiments described herein or perform operations related to the apparatus. In at least some embodiments, such a program is executable by a processor to cause the computer to perform specific operations related to some or all of the blocks of the flowcharts and block diagrams described herein.

[0066] At least some embodiments are described with reference to flowcharts and block diagrams, where the blocks represent (1) steps of a process in which an operation is performed, or (2) sections of a controller responsible for performing an operation. In at least some embodiments, certain steps and sections are implemented by dedicated circuits, programmable circuits supplied with computer-readable instructions stored on a computer-readable medium, and / or processors supplied with computer-readable instructions stored on a computer-readable medium. In at least some embodiments, the dedicated circuits include digital and / or analog hardware circuits, including integrated circuits (ICs) and / or discrete circuits. In at least some embodiments, the programmable circuits include reconfigurable hardware circuits comprising memory elements such as logical AND, OR, XOR, NAND, NOR, and other logical operations, flip-flops, registers, field programmable gate arrays (FPGAs), programmable logic arrays (PLAs), etc.

[0067] In at least some embodiments, a computer-readable storage medium includes a tangible device that can hold and store instructions for use by an instruction execution device. In some embodiments, a computer-readable storage medium includes, but is not limited to, for example, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. A non-exhaustive list of specific examples of computer-readable storage media includes portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disks (DVD), memory sticks, floppy disks, mechanically encoded devices such as punch cards or raised structures in grooves having instructions recorded thereon, and any suitable combination thereof. A computer-readable storage medium as used herein should not be construed to be a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.

[0068] In at least some embodiments, the computer-readable program instructions described herein are downloadable from a computer-readable storage medium to respective computing / processing devices or to an external computer or external storage device via a network, such as, for example, the Internet, a local area network, a wide area network, and / or a wireless network. In at least some embodiments, the network includes copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. In at least some embodiments, a network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage on a computer-readable storage medium in each respective computing / processing device.

[0069] In at least some embodiments, the computer-readable program instructions for performing the operations described above are in any combination of one or more programming languages, including assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or object-oriented programming languages such as Smalltalk, C++, and conventional procedural programming languages such as the "C" programming language or similar programming languages, written in either source code or object code. In at least some embodiments, the computer-readable program instructions are executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer, partially on a remote computer, or entirely on a remote computer or server. In at least some embodiments, in the latter scenario, the remote computer is connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or is connected to an external computer (e.g., via the Internet using an Internet service provider). In at least some embodiments, for example, an electronic circuit including a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) executes the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to customize the electronic circuit for performing aspects of the present disclosure.

[0070] Although embodiments of the present disclosure have been described, the technical scope of any of the problems described in the claims is not limited to the above embodiments. Those skilled in the art will understand that various changes and improvements to the above embodiments are possible. Also, those skilled in the art will understand from the description of the claims that embodiments with such changes or improvements are also included in the technical scope of the present invention.

[0071] The operations, procedures, steps, and stages of each process implemented by the apparatus, system, program, and method shown in the claims, embodiments, or drawings are not explicitly stated as "before" or "preceding", etc., and can be realized in any order unless the output of the previous process is used in the subsequent process. Even if the flow of the process is described using words such as "first" and "next" in the claims, embodiments, or drawings, it does not necessarily mean that the process must be performed in the described order.

[0072] According to at least some embodiments of the present disclosure, a predictively robust model embeds the distribution of each time series dataset among a plurality of time series datasets into a feature vector, predicts a future feature vector of the distribution of a future dataset based on the feature vectors of each time series dataset among the plurality of time series datasets, creates a future dataset from the future feature vector, perturbs the future dataset to generate a plurality of perturbed future datasets, and trains a learning function using the future dataset and each perturbed future dataset to generate a model.

[0073] Some embodiments include instructions in a computer program, a method implemented by a processor that executes the instructions of the computer program, and an apparatus that implements the method. In some embodiments, the apparatus includes a controller that includes a circuit configured to perform the operations within the instructions.

[0074] The above has outlined the features of several embodiments so that those skilled in the art can preferably understand the aspects of the present disclosure. Those skilled in the art should understand that they can easily use the present disclosure as a basis for designing or modifying other processes and structures to perform the same purpose as the embodiments introduced in this specification and / or achieve the same advantages. Those skilled in the art should also understand that such equivalent configurations do not deviate from the spirit and scope of the present disclosure, and various changes, substitutions, and modifications can be made herein without departing from the spirit and scope of the present disclosure.

[0075] Some or all of the above-exemplified embodiments may be described as follows in the appended notes, but are not limited thereto.

[0076] (Appendix 1) Embedding the distribution of each time-series dataset among a plurality of time-series datasets into a feature vector, Predicting a future feature vector of the distribution of a future dataset based on the feature vectors of each time-series dataset among a plurality of time-series datasets, Creating a future dataset from the future feature vector, Perturbing the future dataset to generate a plurality of perturbed future datasets, Training a learning function using the future dataset and each perturbed future dataset to generate a model A computer-readable medium including computer-executable instructions for causing a computer to perform operations including the above.

[0077] (Appendix 2) The computer-readable medium according to Appendix 1, wherein each perturbed future dataset diverges from the future dataset within a predetermined divergence limit.

[0078] (Appendix 3) The computer-readable medium according to Appendix 2, wherein the divergence limit is based on the difference between the future dataset and the latest time-series dataset.

[0079] (Appendix 4) The computer-readable medium according to Appendix 3, wherein the divergence limit is greater than or equal to the difference between a future dataset and the latest time-series dataset.

[0080] (Appendix 5) The computer-readable medium according to Appendix 1, wherein the operation further includes grouping time-series data into a plurality of time-series datasets.

[0081] (Appendix 6) Embedding the distribution includes estimating the density function of each time-series dataset among a plurality of time-series datasets, and embedding the density function of each time-series dataset. The computer-readable medium according to Appendix 1.

[0082] (Appendix 7) The computer-readable medium according to Appendix 1, wherein predicting includes determining a data drift trend.

[0083] (Appendix 8) Predicting includes training a trend estimator to output a subsequent feature vector in time response to an application to each feature vector except the latest feature vector, and applying the trend estimator to the latest feature vector to output a future feature vector. The method according to Appendix 1.

[0084] (Appendix 9) The computer-readable medium according to Appendix 1, wherein creating includes estimating the density function of a future dataset.

[0085] (Appendix 10) The computer-readable medium according to Appendix 1, wherein creating includes generating sample weights based on the density function of a future dataset and the density function of the latest dataset among a plurality of time-series datasets.

[0086] (Appendix 11) Embedding the distribution of each time series dataset among a plurality of time series datasets into a feature vector; Predicting a future feature vector of the distribution of a future dataset based on the feature vectors of each time series dataset among a plurality of time series datasets; Creating a future dataset from the future feature vector; Perturbing the future dataset to generate a plurality of perturbed future datasets; Training a learning function using the future dataset and each perturbed future dataset to generate a model A method comprising:

[0087] (Appendix 12) The method according to Appendix 11, wherein each perturbed future dataset diverges from the future dataset within a predetermined divergence limit.

[0088] (Appendix 13) The method according to Appendix 12, wherein the divergence limit is based on the difference between the future dataset and the latest time series dataset.

[0089] (Appendix 14) The method according to Appendix 13, wherein the divergence limit is greater than or equal to the difference between the future dataset and the latest time series dataset.

[0090] (Appendix 15) The predicting step comprises: Training a trend estimator to output a temporally subsequent feature vector in response to application to each feature vector excluding the latest feature vector; Applying the trend estimator to the latest feature vector to output a future feature vector The method according to Appendix 11, comprising:

[0091] (Appendix 16) Embedding the distribution of each time series dataset among a plurality of time series datasets into a feature vector, Predicting a future feature vector of the distribution of a future dataset based on the feature vectors of each time series dataset among a plurality of time series datasets, Creating a future dataset from the future feature vector, Perturbing the future dataset to generate a plurality of perturbed future datasets, Training a learning function using the future dataset and each perturbed future dataset to generate a model A controller including a circuit configured to perform An apparatus comprising

[0092] (Appendix 17) The apparatus according to Appendix 16, wherein each perturbed future dataset diverges from the future dataset within a predetermined divergence limit. Including, but not limited to, the following appendices.

[0093] (Appendix 18) The apparatus according to Appendix 17, wherein the divergence limit is based on the difference between the future dataset and the latest time series dataset.

[0094] (Appendix 19) The apparatus according to Appendix 18, wherein the divergence limit is greater than or equal to the difference between the future dataset and the latest time series dataset.

[0095] (Appendix 20) The circuit is Training a trend estimator to output a temporally subsequent feature vector in response to application to each feature vector excluding the latest feature vector, Applying the trend estimator to the latest feature vector to output a future feature vector The apparatus according to Appendix 16, further configured to perform

[0096] This application claims the benefit of priority based on U.S. Patent Application No. 17 / 863,338, filed on July 12, 2022, the disclosure of which is incorporated herein by reference in its entirety.

Claims

1. Embedding the distribution of each time series dataset among a plurality of time series datasets into a feature vector; Predicting a future feature vector of the distribution of a future dataset based on the feature vectors of each time series dataset among the plurality of time series datasets; Creating the future dataset from the future feature vector; Perturbing the future dataset to generate a plurality of perturbed future datasets; Training a learning function using the future dataset and each perturbed future dataset to generate a model; A computer-readable medium including instructions executable by a computer for causing the computer to perform operations including the above.

2. The computer-readable medium according to claim 1, wherein each perturbed future dataset diverges from the future dataset within a predetermined divergence limit.

3. The computer-readable medium according to claim 2, wherein the divergence limit is based on the difference between the future dataset and the latest time series dataset.

4. The computer-readable medium according to claim 3, wherein the divergence limit is greater than or equal to the difference between the future dataset and the latest time series dataset.

5. The computer-readable medium according to claim 1, wherein the operations further include grouping time series data into the plurality of time series datasets.

6. The embedding of the distribution includes: estimating a density function of each time series dataset among the plurality of time series datasets; embedding the density function of each time series dataset. The computer-readable medium according to claim 1.

7. The computer-readable medium according to claim 1, wherein the predicting includes determining a data drift trend.

8. The predicting includes: training a trend estimator to output a temporally subsequent feature vector in response to application to each feature vector excluding the latest feature vector; applying the trend estimator to the latest feature vector to output the future feature vector. The computer-readable medium according to claim 1.

9. The computer-readable medium according to claim 1, wherein the creating includes estimating a density function of the future dataset.

10. The creating includes generating sample weights based on the density function of the future data set and the density function of the latest data set among the plurality of time series data sets, the computer-readable medium according to claim 1.

11. Embedding the distribution of each time series data set among the plurality of time series data sets into a feature vector; Predicting a future feature vector of the distribution of a future data set based on the feature vectors of each time series data set among the plurality of time series data sets; Creating the future data set from the future feature vector; Perturbing the future data set to generate a plurality of perturbed future data sets; Training a learning function using the future data set and each perturbed future data set to generate a model A method comprising.

12. The method according to claim 11, wherein each perturbed future data set diverges from the future data set within a predetermined divergence limit.

13. The method according to claim 12, wherein the divergence limit is based on the difference between the future data set and the latest time series data set.

14. The method according to claim 13, wherein the divergence limit is greater than or equal to the difference between the future data set and the latest time series data set.

15. The step of predicting includes: Training a trend estimator to output a temporally subsequent feature vector in response to application to each feature vector excluding the latest feature vector; Applying the trend estimator to the latest feature vector to output the future feature vector The method according to claim 11.

16. Embedding the distribution of each time series data set among the plurality of time series data sets into a feature vector; Predicting a future feature vector of the distribution of a future data set based on the feature vectors of each time series data set among the plurality of time series data sets; Creating the future data set from the future feature vector; Perturbing the future data set to generate a plurality of perturbed future data sets; Training a learning function using the future data set and each perturbed future data set to generate a model A controller including a circuit configured to perform An apparatus comprising.

17. The apparatus according to claim 16, wherein each perturbed future dataset diverges from the future dataset within a predetermined divergence limit. **Claim 18** The apparatus according to claim 17, wherein the divergence limit is based on a difference between the future dataset and the most recent time series dataset. **Claim 19** The apparatus according to claim 18, wherein the divergence limit is greater than or equal to the difference between the future dataset and the most recent time series dataset. **Claim 20** The circuit trains a trend estimator to output a subsequent feature vector in response to application to each feature vector except the most recent feature vector, and applies the trend estimator to the most recent feature vector to output the future feature vector The apparatus according to claim 16, further configured to perform.

Citation Information

Patent Citations

  • Apparatus and method for analyzing time-series data based on machine learning

    US20200380409A1