Accuracy estimation model generation method and accuracy estimation model generation program

WO2026203301A1PCT designated stage Publication Date: 2026-10-01FUJITSU LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/012803
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2026-10-01

Smart Images

  • Figure JP2025012803_01102026_PF_FP_ABST
    Figure JP2025012803_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention improves the precision of accuracy estimation that estimates the prediction accuracy of a prediction model. An information processing device (10) calculates a feature quantity (17a) from time-series data (16) using a certain algorithm. The information processing device (10) acquires a feature quantity (17b), which is calculated from the time-series data (16) using a trained machine learning model (14), by inputting the time-series data (16) to the trained machine learning model (14). By inputting the time-series data (16) to a prediction model (13), the information processing device (10) evaluates prediction accuracy (19) of data for a period later than the time-series data (16), the prediction accuracy (19) being predicted by the prediction model (13). The information processing device (10) generates an accuracy estimation model (15) that has learned the relationship between the prediction accuracy (19) and an input feature quantity (18) based on the feature quantity (17a) and the feature quantity (17b).
Need to check novelty before this filing date? Find Prior Art

Description

Accuracy estimation model generation method and accuracy estimation model generation program

[0001] The present invention relates to an accuracy estimation model generation method and an accuracy estimation model generation program.

[0002] A computer sometimes analyzes time-series data indicating temporal changes of a certain index value. As time-series analysis, a computer may use a prediction model to predict an index value for a period later than the time-series data from certain time-series data.

[0003] A computer can use a plurality of prediction models such as an autoregressive model and a moving average model. These plurality of prediction models have different fields of specialization from each other. The prediction accuracy of each prediction model depends on properties of the time-series data such as seasonal fluctuations and external factors. Therefore, a computer sometimes uses different prediction models depending on time-series data.

[0004] For example, a Feature-based Forecast Model Selection (FFORMS) system that selects a statistical model suitable for input time-series data from among a plurality of statistical models has been proposed. The FFORMS system inputs samples of time-series data to each of a plurality of statistical models, and determines an optimal statistical model having the highest prediction accuracy for the input samples. The FFORMS system also calculates statistical feature quantities such as trend intensity and seasonal intensity from the samples. The FFORMS system learns the relationship between the statistical feature quantities and the optimal statistical model. The FFORMS system calculates a statistical feature quantity from input time-series data, and estimates an optimal statistical model from the statistical feature quantity in accordance with a learning result.

[0005] Thiyanga S. Talagala, Rob J. Hyndman, and George Athanasopoulos, "Meta-learning How to Forecast Time Series", Journal of Forecasting, Volume 42, Issue 6, pages 1476-1501, February 9, 2023

[0006] The aforementioned FFORMS system uses a statistical model as its predictive model and statistical features as features of time series data. However, among the various predictive models available to computers, other types of predictive models, such as neural networks, may be included. Depending on the type of predictive model, it can be difficult for a computer to accurately estimate prediction accuracy from statistical features alone. Therefore, in one aspect, the present invention aims to improve the accuracy of accuracy estimation for estimating the prediction accuracy of a predictive model.

[0007] One aspect of this method provides a computer-based method for generating an accuracy estimation model, which involves calculating a first feature from time-series data using a specific algorithm, inputting the time-series data into a trained machine learning model to obtain a second feature calculated from the time-series data using the trained machine learning model, inputting the time-series data into a prediction model to evaluate the prediction accuracy of data for periods later than the time-series data, and generating an accuracy estimation model that learns the relationship between input features based on the first and second features and the prediction accuracy.

[0008] In one aspect, the accuracy of the accuracy estimate used to estimate the prediction accuracy of the prediction model is improved. The above and other objects, features and advantages of the present invention will become apparent from the following description in conjunction with the accompanying drawings illustrating preferred embodiments as examples of the present invention.

[0009] This is a diagram illustrating the information processing device of the first embodiment. This is a diagram showing an example of the hardware of the information processing device of the second embodiment. This is a diagram showing an example of time series data. This is a diagram showing an example of estimating prediction accuracy using an accuracy estimation model. This is a diagram showing an example of generating input features using a base model and an association estimation model. This is a diagram showing an example of the structure of a base model. This is a diagram showing an example of converting an intermediate vector set to an aggregate vector. This is a diagram showing an example of statistical feature components. This is a diagram showing an example of generating additional features using an association estimation model. This is a block diagram showing an example of the functions of the information processing device. This is a flowchart showing an example of a machine learning procedure. This is a flowchart showing an example of a machine learning procedure (continued 1). This is a flowchart showing an example of a machine learning procedure (continued 2). This is a flowchart showing an example of a time series forecast procedure. This is a graph showing a comparative example of accuracy estimation.

[0010] Hereinafter, this embodiment will be described with reference to the drawings. (a) First Embodiment Figure 1 is a diagram illustrating the information processing device of the first embodiment. The information processing device 10 of the first embodiment generates an accuracy estimation model by machine learning. The accuracy estimation model estimates the accuracy of a prediction model that predicts time series data. The information processing device 10 may be a client device or a server device. The information processing device 10 may also be called a computer, a machine learning device, or an accuracy estimation model generation device.

[0011] The information processing device 10 includes a storage unit 11 and a processing unit 12. The storage unit 11 may be a volatile memory such as Random Access Memory (RAM). Alternatively, the storage unit 11 may be a non-volatile storage such as a Hard Disk Drive (HDD) or Solid State Drive (SSD).

[0012] The processing unit 12 is a processor, such as a Central Processing Unit (CPU), Graphics Processing Unit (GPU), or Digital Signal Processor (DSP). However, the processing unit 12 may also include electronic circuits such as an Application Specific Integrated Circuit (ASIC) or a Field Programmable Gate Array (FPGA). The processor executes a program stored in memory, such as RAM. The processor is sometimes called a processor circuit. A collection of processors is sometimes called a multiprocessor or simply a "processor." Different processes among the multiple processes described later may be executed by different processors.

[0013] The memory unit 11 stores time-series data 16. The time-series data 16 shows the time change of a certain index value. The index value is, for example, the demand for a certain product. The time-series data 16 contains multiple values ​​corresponding to multiple times. The multiple times have fixed intervals, such as hourly intervals, daily intervals, monthly intervals, or yearly intervals. The period shown by the time-series data 16 may be called the input period, past period, or training period.

[0014] The memory unit 11 also stores the prediction model 13. The prediction model 13 predicts data for a period later than the input time series data from the input time series data. The prediction model 13 may be a statistical model such as an autoregressive model. Alternatively, the prediction model 13 may be a machine learning model generated from training data by machine learning, or it may be a neural network such as a deep neural network. The predicted data includes one or more values ​​corresponding to one or more time points. The predicted data may be time series data. The later period may be a subsequent period following the input time series data. The later period may be called the prediction period, future period, or test period.

[0015] Furthermore, the memory unit 11 stores the machine learning model 14. The machine learning model 14 has been trained by machine learning using training data. The machine learning model 14 calculates features from the input time series data. The calculated features may be vector data containing multiple dimensions. The machine learning model 14 is, for example, a neural network such as a transformer. The machine learning model 14 may be trained using large-scale training data containing various types of time series data, and may be a large-scale machine learning model containing many layers and many parameters. The machine learning model 14 may also be called the base model.

[0016] The processing unit 12 calculates a feature vector 17a (first feature vector) from the time series data 16 using a specific algorithm. The specific algorithm is, for example, a statistical algorithm. The feature vector 17a may include statistics such as the first-order autocorrelation coefficient, the first-order autocorrelation coefficient of the difference time series data, the first-order autocorrelation coefficient of the second-order difference time series data, trend strength, or seasonality strength. The feature vector 17a may also be vector data containing two or more statistics.

[0017] Furthermore, the processing unit 12 obtains feature vectors 17b (second feature vectors) by inputting the time-series data 16 into the machine learning model 14. Feature vectors 17b are calculated from the time-series data 16 using the machine learning model 14. Feature vectors 17b may be the feature vectors themselves calculated by the machine learning model 14, or they may be generated by transforming the calculated feature vectors.

[0018] For example, the processing unit 12 extracts features from an intermediate layer among the multiple layers included in the machine learning model 14. The extracted features may be vector data. The vector data may include multiple vectors corresponding to multiple time points included in the time series data 16. The processing unit 12 may generate feature 17b by aggregating the multiple vectors. Feature 17b may include statistics such as the maximum value, mean, or variance of the multiple vectors.

[0019] Furthermore, the processing unit 12 evaluates the prediction accuracy 19 corresponding to the time series data 16 by inputting the time series data 16 into the prediction model 13. The prediction accuracy 19 is the prediction accuracy of data for a period later than the time series data 16, as predicted by the prediction model 13. The indicator for prediction accuracy 19 may be a prediction error, where a smaller value indicates higher prediction accuracy, or it may be the Mean Absolute Scaled Error (MASE). Alternatively, the indicator for prediction accuracy 19 may be a flag indicating the quality of the prediction by the prediction model 13, or it may be the ranking of prediction accuracy among multiple prediction models.

[0020] For example, the processing unit 12 acquires ground truth data showing actual results for a period after the time series data 16. The ground truth data is stored in the storage unit 11, for example, in association with the time series data 16. The processing unit 12 also inputs the time series data 16 into the prediction model 13 to acquire predicted data output by the prediction model 13. The processing unit 12 calculates the prediction accuracy 19 by comparing the predicted data with the ground truth data. The processing unit 12 may change the execution order of calculating the feature quantity 17a, acquiring the feature quantity 17b, and evaluating the prediction accuracy 19.

[0021] The processing unit 12 uses feature quantities 17a and 17b and prediction accuracy 19 to generate an accuracy estimation model 15 that has learned the relationship between the input feature quantity 18 and the prediction accuracy 19. For example, the processing unit 12 generates training data from multiple samples corresponding to time series data 16, which includes multiple records corresponding to pairs of input feature quantity 18 and prediction accuracy 19. The processing unit 12 uses this training data to train the accuracy estimation model 15.

[0022] The accuracy estimation model 15 can estimate the prediction accuracy of the prediction model 13 for time series data containing the received input features, without running the prediction model 13. The accuracy estimation model 15 may also estimate the prediction accuracy of multiple prediction models, including the prediction model 13. In that case, the accuracy estimation model 15 learns the relationship between the input features 18 and the multiple prediction accuracies corresponding to the multiple prediction models.

[0023] The input feature 18 is based on feature 17a and feature 17b. The input feature 18 may include feature 17a and feature 17b. Alternatively, the input feature 18 may include feature 17a and a feature converted from feature 17b (a third feature).

[0024] For example, the processing unit 12 calculates a filtered feature by removing the values ​​of some of the dimensions from the feature quantity 17b. The third feature may be the filtered feature. Alternatively, the processing unit 12 may perform dimensionality reduction on either the feature quantity 17b or the filtered feature. Dimensionality reduction may be performed using Principal Component Analysis (PCA) or an autoencoder. The third feature may be the feature after dimensionality reduction. As a result, the third feature has fewer dimensions than the feature quantity 17b.

[0025] The processing unit 12 may use feature quantities 17a and 17b to generate an association estimation model that learns the relationships between multiple dimensions contained in feature quantities 17a and 17b. The processing unit 12 may use the association estimation model to determine which dimensions to remove. For example, the processing unit 12 has the association estimation model estimate the values ​​of each of the multiple dimensions contained in feature quantity 17b from feature quantity 17a. The processing unit 12 then determines which dimensions to remove based on the estimation accuracy of the association estimation model for the multiple dimensions.

[0026] The processing unit 12 may prioritize selecting dimensions with high estimation accuracy from among multiple dimensions as some of the dimensions mentioned above. Since the values ​​of dimensions with high estimation accuracy can be estimated from feature quantity 17a, they represent features similar to feature quantity 17a, and therefore the benefit of using them in combination with feature quantity 17a is small. On the other hand, since the values ​​of dimensions with low estimation accuracy are difficult to estimate from feature quantity 17a, they represent features not similar to feature quantity 17a, and therefore the benefit of using them in combination with feature quantity 17a is large. In this way, the processing unit 12 can efficiently reduce the size of the input feature quantity 18.

[0027] As described above, the information processing device 10 of the first embodiment calculates feature quantities 17a from time series data 16 using a certain algorithm. The information processing device 10 inputs the time series data 16 into a trained machine learning model 14 to obtain feature quantities 17b calculated from the time series data 16 using the trained machine learning model 14. The information processing device 10 inputs the time series data 16 into a prediction model 13 to evaluate the prediction accuracy 19 of data for a period later than the time series data 16, which is predicted by the prediction model 13. The information processing device 10 generates an accuracy estimation model 15 that has learned the relationship between the input feature quantities 18 based on feature quantities 17a and 17b and the prediction accuracy 19.

[0028] As a result, the information processing device 10 can use the accuracy estimation model 15 to estimate the prediction accuracy of the prediction model 13 for given time series data without actually running the prediction model 13. Therefore, the information processing device 10 can efficiently determine whether or not the prediction model 13 is useful for given time series data. For example, the information processing device 10 can efficiently use different prediction models depending on the time series data.

[0029] Furthermore, the information processing device 10 uses not only algorithm-based features 17a such as statistics, but also features 17b calculated using the machine learning model 14. Features 17b complement the features of the time series data 16 that are not represented by features 17a. Therefore, the accuracy estimation model 15 can estimate the prediction accuracy of the prediction model 13 with high accuracy, regardless of the type of prediction model 13. In particular, when the prediction model 13 is a machine learning model such as a neural network, the accuracy estimation model 15 can estimate the prediction accuracy with high accuracy because of the high affinity between the prediction model 13 and the machine learning model 14.

[0030] (b) Figure 2 of the second embodiment shows an example of the hardware of the information processing device of the second embodiment. The information processing device 100 of the second embodiment predicts future time series data following input time series data using a trained predictive model. The information processing device 100 dynamically selects one or more predictive models from a plurality of predictive models that are suitable for the input time series data.

[0031] The information processing device 100 uses a trained accuracy estimation model to estimate the prediction accuracy of each prediction model for the input time series data before executing each prediction model. The information processing device 100 trains the accuracy estimation model using machine learning. Note that the training of the accuracy estimation model, accuracy estimation, and prediction of time series data may be performed by different information processing devices. The information processing device 100 corresponds to the information processing device 10 of the first embodiment.

[0032] The information processing device 100 includes a CPU 101, RAM 102, HDD 103, GPU 104, input interface 105, media reader 106, and communication interface 107. The CPU 101 corresponds to the processing unit 12 of the first embodiment. The RAM 102 or HDD 103 corresponds to the storage unit 11 of the first embodiment.

[0033] The CPU 101 is a processor that executes program instructions. The CPU 101 loads the program and data from the HDD 103 into the RAM 102 and executes the program. The information processing device 100 may have multiple processors.

[0034] RAM 102 is a volatile semiconductor memory that temporarily stores programs executed by the CPU 101 and data used for calculations by the CPU 101. The information processing device 100 may have a type of volatile memory other than RAM.

[0035] The HDD 103 is a non-volatile storage device that stores software programs such as the operating system (OS), middleware, and application software, as well as data. The information processing device 100 may also have other types of non-volatile storage, such as an SSD or flash memory.

[0036] The GPU 104 works in conjunction with the CPU 101 to perform image processing and outputs the image to the display device 111 connected to the information processing device 100. The display device 111 is, for example, a cathode ray tube (CRT) display, a liquid crystal display, an organic EL (Electro Luminescence) display, or a projector.

[0037] Furthermore, the GPU 104 may be used as a General Purpose Computing on Graphics Processing Unit (GPGPU). The GPU 104 may execute a program in response to instructions from the CPU 101. The information processing device 100 may have a volatile semiconductor memory other than RAM 102 as GPU memory.

[0038] The input interface 105 receives input signals from an input device 112 connected to the information processing device 100. The input device 112 is, for example, a mouse, a touch panel, or a keyboard. Multiple input devices may be connected to the information processing device 100.

[0039] The media reader 106 is a reading device that reads programs and data recorded on a recording medium 113. The recording medium 113 is, for example, a magnetic disk, an optical disk, or a semiconductor memory. Magnetic disks include flexible disks (FDs) and HDDs. Optical disks include Compact Discs (CDs) and Digital Versatile Discs (DVDs). The media reader 106 copies the programs and data read from the recording medium 113 to another recording medium such as the RAM 102 or the HDD 103. The read program may be executed by the CPU 101.

[0040] The recording medium 113 may be a portable recording medium. The recording medium 113 may be used for distribution of programs and data. Furthermore, the recording medium 113 and the HDD 103 may be referred to as computer-readable recording media.

[0041] The communication interface 107 communicates with other information processing apparatuses via a network 114. The communication interface 107 may be a wired communication interface connected to a wired communication apparatus such as a switch or a router. Alternatively, the communication interface 107 may be a wireless communication interface connected to a wireless communication apparatus such as a base station or an access point. Next, prediction of time-series data will be described.

[0042] FIG. 3 is a diagram illustrating an example of time-series data. A table 141 includes the time-series data. In the example of FIG. 3, the time-series data indicates temporal changes in monthly product sales volume. A graph 142 is a graph that visualizes the time-series data included in the table 141.

[0043] The table 141 includes time-series data for an input period. The input period is a period to which the time-series data input to a prediction model belongs, and may be referred to as a past period or a training period. In the example of FIG. 3, the time-series data for the input period indicates product sales volume for seven months.

[0044] Further, table 141 includes four types of time-series data for a prediction period predicted by prediction models A, B, C, and D. The prediction period is a period to which the time-series data output from a prediction model belongs, and is sometimes referred to as a future period or a test period. Normally, the prediction period is a subsequent period consecutive to the input period. In the example of FIG. 3, the time-series data of the prediction period indicates the number of product units sold over five months following the input period.

[0045] As graph 142 shows, a plurality of prediction models may output different time-series data for the prediction period from the same time-series data related to the input period. The prediction accuracy of each prediction model depends on properties inherent to the time-series data, such as seasonal fluctuations and external factors. Fields of time-series data for which prediction accuracy is high differ depending on the prediction model. Accordingly, the information processing apparatus 100 executes ensemble prediction that uses a plurality of prediction models in combination.

[0046] For example, the information processing apparatus 100 outputs the average of prediction results obtained from a plurality of prediction models as a final prediction result. Further, for example, the information processing apparatus 100 causes each of the plurality of prediction models to predict a numerical value at the latest point in time for which the correct answer is already known. The information processing apparatus 100 uses one or a small number of prediction models, among the plurality of prediction models, that output a numerical value having a small error from the correct answer. Further, for example, the information processing apparatus 100 calculates a weighted average obtained by weighting prediction results of the plurality of prediction models such that a prediction model with a smaller error is assigned a larger weight.

[0047] However, executing all prediction models for each time-series prediction consumes a large amount of computational resources and a long computation time. Accordingly, before executing each prediction model, the information processing apparatus 100 estimates a prediction model that is useful for the current time-series data among the plurality of prediction models. A useful prediction model is one or more prediction models having high prediction accuracy.

[0048] The information processing device 100 restricts the prediction models to be executed according to the estimation results. The information processing device 100 may select prediction models whose prediction accuracy exceeds a threshold. Alternatively, the information processing device 100 may select a certain number of prediction models with relatively high prediction accuracy. For the estimation of useful prediction models, the information processing device 100 generates an accuracy estimation model that estimates the prediction accuracy of each of the multiple prediction models from the features of the time series data of the input period.

[0049] Figure 4 shows an example of estimating prediction accuracy using an accuracy estimation model. In the training phase, the information processing device 100 trains the accuracy estimation model 132. In the operation phase, the information processing device 100 estimates prediction accuracy using the accuracy estimation model 132. For simplicity of explanation, in the example in Figure 4, the accuracy estimation model 132 estimates the prediction accuracy of two prediction models, but it is possible to estimate the prediction accuracy of three or more prediction models.

[0050] During the training phase, the information processing device 100 acquires time series data 143a and 143b. Time series data 143a is the time series data for the input period. Time series data 143b is the correct time series data for the prediction period. The information processing device 100 may generate time series data 143a and 143b by dividing a sample of time series data into two periods.

[0051] The information processing device 100 inputs time series data 143a to prediction model 131a and prediction model 131b, respectively. Prediction model 131a or prediction model 131b may be a statistical model or a machine learning model such as a neural network. Prediction model 131a outputs time series data 144a. Prediction model 131b outputs time series data 144b. Time series data 144a and 144b are prediction results of time series data for the prediction period, respectively.

[0052] The information processing device 100 calculates the prediction accuracy 145a of the prediction model 131a for time series data 143a by comparing time series data 143b and time series data 144a. The information processing device 100 also calculates the prediction accuracy 145b of the prediction model 131b for time series data 143a by comparing time series data 143b and time series data 144b. The prediction accuracy values ​​145a and 145b are, for example, MASE. However, the information processing device 100 may use a prediction accuracy metric other than MASE.

[0053] Furthermore, the information processing device 100 generates input features 146 from time series data 143a. The input features 146 may also be called time series features. The information processing device 100 trains the accuracy estimation model 132 using the prediction accuracies 145a, 145b and the input features 146. The input features 146 correspond to the explanatory variables of the accuracy estimation model 132. The multiple prediction accuracies, including prediction accuracies 145a, 145b, correspond to the target variables of the accuracy estimation model 132.

[0054] The accuracy estimation model 132 is a machine learning model. The accuracy estimation model 132 may be a deep machine learning model using a neural network, or it may be another type of machine learning model such as a decision tree. The accuracy estimation model 132 may also estimate whether the prediction accuracy of each of the multiple prediction models exceeds a threshold. The threshold may be the prediction accuracy of a specific prediction model. Alternatively, the accuracy estimation model 132 may estimate the prediction model with the highest prediction accuracy among the multiple prediction models. Alternatively, the accuracy estimation model 132 may estimate a certain number of prediction models in descending order of prediction accuracy.

[0055] During the operational phase, the information processing device 100 acquires time-series data 147. Time-series data 147 is time-series data for the input period. The information processing device 100 generates input features 148 from the time-series data 147 in the same way as input features 146. The information processing device 100 inputs the input features 148 into the accuracy estimation model 132. The accuracy estimation model 132 estimates prediction accuracy 149a and prediction accuracy 149b. Prediction accuracy 149a is an estimated value of the prediction accuracy of prediction model 131a for the time-series data 147. Prediction accuracy 149b is an estimated value of the prediction accuracy of prediction model 131b for the time-series data 147.

[0056] The information processing device 100 uses the prediction accuracies 149a and 149b to decide whether or not to use prediction models 131a and 131b, respectively. For example, the information processing device 100 decides to use prediction model 131a and not prediction model 131b. In that case, the information processing device 100 inputs the time series data 147 into prediction model 131a and does not input the time series data 147 into prediction model 131b. The information processing device 100 uses the output of prediction model 131a to generate time series data for the prediction period following the time series data 147. Next, the input features input to the accuracy estimation model 132 will be described.

[0057] Figure 5 shows an example of generating input features using a base model and an association estimation model. The information processing device 100 acquires time series data 151. The time series data 151 is time series data for the input period. The information processing device 100 calculates statistical features 152 from the time series data 151 using a certain algorithm. The statistical features 152 is a numerical vector containing multiple statistics related to the time series data 151.

[0058] Furthermore, the information processing device 100 inputs time-series data 151 to the base model 133. The base model 133 is a deep machine learning model with numerous parameters, trained using a large amount of training data. As will be described later, the base model 133 is implemented, for example, using a transformer. The information processing device 100 may use the machine learning model described in the following non-patent document as the base model 133.

[0059] Abdul Fatir Ansari, Lorenzo Stella, Caner Turkmen, Xiyuan Zhang, Pedro Mercado, Huibin Shen, Oleksandr Shchur, Syama Sundar Rangapuram, Sebastian Pineda Arango, Shubham Kapoor, Jasper Zschiegner, Danielle C. Maddix, Hao Wang, Michael W. Mahoney, Kari Torkkola, Andrew Gordon Wilson, Michael Bohlke-Schneider, and Yuyang Wang, "Chronos: Learning the Language of Time Series", Transactions on Machine Learning Research, November 12, 2024.

[0060] The information processing device 100 extracts an intermediate vector 153 from an intermediate layer among the multiple layers included in the base model 133. The number of dimensions of the intermediate vector 153 depends on which layer it is extracted from. For example, the intermediate vector 153 is a 768-dimensional numerical vector. The base model 133 calculates the intermediate vector 153 from each of the multiple time values ​​included in the time series data 151. Therefore, the information processing device 100 extracts multiple intermediate vectors corresponding to the intermediate vector 153 from the base model 133.

[0061] The information processing device 100 converts multiple intermediate vectors corresponding to multiple time points into an aggregate vector 154. For example, the information processing device 100 calculates statistics such as the average of the values ​​of the multiple intermediate vectors for each dimension. The aggregate vector 154 may be a numerical vector that enumerates statistics for multiple dimensions. The number of dimensions of the aggregate vector 154 may be the same as that of the intermediate vector 153, or it may be an integer multiple of the number of dimensions of the intermediate vector 153. The number of dimensions of the aggregate vector 154 does not depend on the length of the time series data 151.

[0062] The information processing device 100 uses the learning results of the association estimation model 134 to convert the aggregate vector 154 into additional features 155. The additional features 155 are numerical vectors with fewer dimensions than the aggregate vector 154. The association estimation model 134 learns the relationships between the statistical features 152 and the multiple dimensions contained in the aggregate vector 154.

[0063] The association estimation model 134 estimates the values ​​of each of the multiple dimensions included in the aggregate vector 154 from the statistical features 152. The association estimation model 134 may include multiple multivariate regression models corresponding to the multiple dimensions. The information processing device 100 has previously evaluated the estimation error of the association estimation model 134 for each of the multiple dimensions. The association estimation model 134 classifies the multiple dimensions into dimensions with large estimation errors and dimensions with small estimation errors.

[0064] Values ​​of dimensions with small estimation errors are easy to estimate from statistical feature 152 and exhibit similar characteristics to statistical feature 152. Therefore, there is little benefit in using statistical feature 152 in conjunction with values ​​of dimensions with small estimation errors. On the other hand, values ​​of dimensions with large estimation errors are difficult to estimate from statistical feature 152 and exhibit characteristics that are not similar to statistical feature 152. Therefore, there is a great benefit in using statistical feature 152 in conjunction with values ​​of dimensions with large estimation errors.

[0065] The information processing device 100 extracts values ​​for dimensions with large estimation errors from the aggregate vector 154 and removes values ​​for dimensions with small estimation errors. This allows the information processing device 100 to generate a filtered aggregate vector. The information processing device 100 then generates additional features 155 by performing dimensionality reduction on the filtered aggregate vector. Dimensionality reduction can be performed using, for example, principal component analysis or an autoencoder. The information processing device 100 then generates input features 156 by concatenating the statistical features 152 and the additional features 155.

[0066] Figure 6 shows an example of the structure of a base model. The base model 133 includes a token generation unit 135, encoder blocks 136-1, 136-2, ..., 136-n, and a decoder block 137. The token generation unit 135, encoder blocks 136-1, 136-2, ..., 136-n, and decoder block 137 are connected in series. The base model 133 may include multiple decoder blocks.

[0067] The base model 133 receives time series data 157. The time series data 157 is time series data for the input period. The token generation unit 135 converts the time series data 157 into a token sequence 158. Each token in the token sequence 158 corresponds to a numerical value for a single time in the time series data 157.

[0068] The token generation unit 135 converts each numerical value contained in the time series data 157 into a numerical value within a certain range by scaling. For example, the token generation unit 135 searches for the maximum and minimum values ​​and determines a scaling method such that the maximum value corresponds to 1 and the minimum value corresponds to -1. Next, the token generation unit 135 converts the scaled continuous values ​​into discrete values ​​of a certain number of dimensions by quantization. For example, each token is represented in 768 dimensions.

[0069] The token generation unit 135 may also perform position encoding on the quantized tokens. For example, position encoding calculates a position number for each token indicating its position from the beginning, and converts the position number into a position vector using a sine function and a cosine function. The dimensionality of the position vector is the same as the dimensionality of the original token. Position encoding adds the position vector to the original token.

[0070] Encoder block 136-1 receives the token sequence 158. Encoder block 136-1 converts the received token sequence into another token sequence and outputs the converted token sequence to encoder block 136-2. Encoder blocks 136-2 to 136-n also convert the token sequence. Encoder block 136-n outputs the converted token sequence to decoder block 137.

[0071] Encoder block 136-1 includes a self-attention layer 138 and a feedforward layer 139. Encoder blocks 136-2 to 136-n may have the same structure as encoder block 136-1. The self-attention layer 138 converts the received token sequence using an attention mechanism.

[0072] The self-attention layer 138 has a query matrix, a key matrix, and a value matrix as parameters. The coefficients in these three matrices are trained through machine learning. The self-attention layer 138 selects one token of interest from the token sequence. The self-attention layer 138 transforms the token of interest using the query matrix and calculates a vector called the query. The self-attention layer 138 also transforms each of the multiple tokens using the key matrix and calculates a vector called the key. The self-attention layer 138 calculates the dot product of the query and the key as the attention score for each token. The attention score indicates the degree of relevance between the token of interest and each of the other tokens.

[0073] The self-attention layer 138 transforms each of the multiple tokens using a value matrix and calculates a vector called the value. The self-attention layer 138 uses the attention score as a weight to calculate a weighted sum of the values ​​among the multiple tokens and outputs the calculated weighted sum as the transformed vector for the token of interest. The self-attention layer 138 repeats the above transformation process while changing the token of interest.

[0074] The feedforward layer 139 is a forward-direction neural network. The feedforward layer 139 receives a token sequence from the self-attention layer 138. The feedforward layer 139 uses parameters to individually transform multiple tokens contained in the token sequence. These parameters are trained through machine learning.

[0075] The decoder block 137 receives a token sequence from the encoder block 136-n and converts the received token sequence into a token sequence 159. The length of the token sequence 159 is the same as that of the token sequence 158. Each token in the token sequence 159 corresponds to a single token in the token sequence 158. The processing of the decoder block 137 and the meaning of the token sequence 159 depend on the task to be performed by the base model 133. The decoder block 137 may include a self-attention layer and a feedforward layer.

[0076] The base model 133 may be trained by the information processing device 100, or it may be trained by another information processing device. In the latter case, the information processing device 100 obtains the trained base model 133 from the other information processing device.

[0077] The information processing device 100 trains the base model 133 using, for example, the first or second training method described below. In the first training method, the information processing device 100 designates a delayed token sequence, obtained by shifting multiple tokens in the token sequence 158 one by one backward, as the correct answer for the token sequence 159. The first training method causes the decoder block 137 to perform the task of predicting the value at the next time point from the value at a given time point. The information processing device 100 trains the base model 133 using backpropagation so that the error between the predicted token sequence and the correct token sequence becomes small.

[0078] In the second training method, the information processing device 100 randomly selects some tokens from the token sequence 158 and masks the selected tokens. The information processing device 100 also designates the original token sequence 158 as the correct answer for the token sequence 159. The second training method involves having the decoder block 137 perform the task of predicting missing tokens from other tokens. The information processing device 100 trains the base model 133 using backpropagation so that the error between the predicted token sequence and the correct token sequence is minimized.

[0079] As described above, the information processing device 100 inputs time-series data 151 to the trained base model 133. The information processing device 100 extracts the output of the encoder block 136-n, which is the final stage encoder block, as an intermediate vector 153. At this time, the information processing device 100 extracts a number of intermediate vectors corresponding to the length of the time-series data 151.

[0080] Figure 7 shows an example of conversion from a set of intermediate vectors to an aggregated vector. The information processing device 100 extracts a set of intermediate vectors from the base model 133, such as intermediate vectors 161-1, 161-2, 161-3, 161-4, 161-5, ..., 161-n. The information processing device 100 converts the extracted set of intermediate vectors into an aggregated vector 162 that is independent of the length of the time series data.

[0081] The information processing device 100 extracts numerical values ​​of the same dimension from multiple intermediate vectors and calculates one or more statistical quantities from the extracted numerical values. These statistical quantities may be, for example, the maximum value, mean, or variance. However, the information processing device 100 may also calculate some types of statistical quantities using the algorithm for calculating the aforementioned statistical features 152. As will be described later, an example of an algorithm-based statistical quantity is trend intensity or seasonality intensity.

[0082] If two or more statistical quantities are calculated for a single dimension, the information processing device 100 concatenates these two or more statistical quantities in the aggregated vector 162. The information processing device 100 performs the above process for all dimensions of the intermediate vector. Therefore, the number of dimensions of the aggregated vector 162 is the same as that of the intermediate vector, or an integer multiple of the number of dimensions of the intermediate vector.

[0083] For example, the information processing device 100 calculates the maximum value, mean, and variance for each dimension of the intermediate vector. In this case, the first to third dimensions of the aggregated vector 162 are the maximum value, mean, and variance of the first dimension of the intermediate vector. Also, the fourth to sixth dimensions of the aggregated vector 162 are the maximum value, mean, and variance of the second dimension of the intermediate vector. If the intermediate vector has 768 dimensions, the aggregated vector 162 has 2304 dimensions.

[0084] Figure 8 shows an example of the components of a statistical feature. The information processing device 100 calculates a statistical feature 163 separately from the aggregate vector from the same time series data. The statistical feature 163 includes multiple statistics calculated by a statistical algorithm. In the example in Figure 8, the statistical feature 163 is a vector containing 42 statistics.

[0085] The first dimension of statistical feature 163 is the first-order autocorrelation coefficient of the time series data. Equation (1) shows the k-th order autocorrelation coefficient. The k-th order autocorrelation coefficient is the correlation coefficient between the time series data and the delayed time series data obtained by shifting the numerical values ​​of multiple time points included in the time series data by k. In equation (1), r k is the k-th order autocorrelation coefficient, y t is the numerical value of time t included in the time series data, k is the delay amount, T is the final time of the time series data, y - (above y) - The variables marked with an asterisk (*) are the average of the numerical values ​​included in the time series data. The linear autocorrelation coefficient is calculated by substituting 1 for k in equation (1).

[0086]

[0087] The second dimension of statistical feature 163 is the linear autocorrelation coefficient of the difference time series data. The difference time series data is the difference between the time series data and the delayed time series data obtained by shifting the numerical values ​​of multiple time points included in the time series data by 1. The linear autocorrelation coefficient of the difference time series data is calculated by replacing y in equation (1) with the difference time series data.

[0088] The third dimension of statistical feature 163 is the linear autocorrelation coefficient of the second-order difference time series data. The second-order difference time series data is the difference between the difference time series data and the delayed difference time series data obtained by shifting the numerical values ​​of multiple time points included in the difference time series data by 1. The linear autocorrelation coefficient of the second-order difference time series data is calculated by replacing y in formula (1) with the second-order difference time series data. In the second embodiment, the delay amount for both the difference time series data and the second-order difference time series data is 1, but the delay amount may be 2 or more.

[0089] The fourth dimension of statistical feature 163 is trend intensity. The fifth dimension of statistical feature 163 is seasonality intensity. Trend intensity and seasonality intensity are calculated using Seasonal and Trend Decomposition using Loess (STL) analysis.

[0090] Generally, time series data contains trend components, seasonal components, and noise components. Seasonal components are fluctuation patterns that appear repeatedly at regular intervals, such as daily, weekly, monthly, or yearly cycles. Trend components are long-term fluctuation trends that appear when the seasonal components are removed from time series data. Examples of long-term fluctuation trends include linear monotonic increase, linear monotonic decrease, curvilinear monotonic increase, or curvilinear monotonic decrease. Noise components represent exceptional numerical fluctuations that cannot be explained by either the trend or seasonal components.

[0091] STL analysis decomposes time series data into trend, seasonal, and noise components. For example, STL analysis calculates the trend component from the entire time series data and removes the calculated trend component from the time series data. STL analysis divides the time series data after trend removal into unit periods such as one day, one week, one month, and one year, and calculates the average value for each unit period. STL analysis discovers the periodicity of the seasonal component by searching for the length of the unit period in which the calculated average approaches zero. STL analysis calculates the noise component by further removing the seasonal component from the time series data after trend removal.

[0092] Formula (2) shows the trend intensity. Formula (3) shows the seasonality intensity. In formulas (2) and (3), f t is the trend component at time t, s t is the seasonal component of time t, e t represents the noise component at time t, and Var represents the variance.

[0093]

[0094]

[0095] Trend strength indicates the strength of the trend component relative to the sum of the trend component and the noise component. Trend strength is calculated by dividing the variance of the noise component by the variance of the sum of the trend component and the noise component, and subtracting the quotient from 1. Seasonality strength indicates the strength of the seasonal component relative to the sum of the seasonal component and the noise component. Seasonality strength is calculated by dividing the variance of the noise component by the variance of the sum of the seasonal component and the noise component, and subtracting the quotient from 1.

[0096] The 42nd dimension of the statistical feature 163 is the zero ratio. The zero ratio is the proportion of times in the time series data where the value is 0. The information processing device 100 may change the order of the multiple statistics included in the statistical feature 163.

[0097] Figure 9 shows an example of generating additional features using an association estimation model. In the training phase, the information processing device 100 trains the association estimation model 134. In the test phase, the information processing device 100 generates an effective dimension list 176 using the trained association estimation model 134. In the operation phase, the information processing device 100 reduces the dimension of the aggregation vector using the generated effective dimension list 176.

[0098] During the training phase, the information processing device 100 acquires statistical features 171 and aggregation vectors 172. The statistical features 171 and aggregation vectors 172 are calculated from the same time-series data using the method described above. The information processing device 100 uses the statistical features 171 as input data corresponding to explanatory variables and the aggregation vectors 172 as ground truth labels corresponding to the target variable to train the association estimation model 134.

[0099] For example, the association estimation model 134 includes multiple multivariate regression models corresponding to multiple dimensions of the aggregate vector 172. One multivariate regression model estimates the value of one dimension in the aggregate vector 172 from multiple statistics included in the statistical features 171. In the training phase, the multivariate regression models learn the relationship between multiple statistics and the value of one dimension.

[0100] During the test phase, the information processing device 100 acquires statistical features 173 and aggregate vectors 174. The statistical features 173 and aggregate vectors 174 are calculated from the same time-series data using the method described above. Typically, the information processing device 100 uses different time-series data for the training phase and the test phase.

[0101] The information processing device 100 uses statistical features 173 as input data corresponding to explanatory variables and aggregate vectors 174 as ground truth labels corresponding to the target variable to evaluate the estimation accuracy of the association estimation model 134. The information processing device 100 inputs the statistical features 173 into the association estimation model 134 and obtains the estimated aggregate vectors 175 from the association estimation model 134. The information processing device 100 compares aggregate vectors 174 and 175 and calculates the estimation error for each of the multiple dimensions included in aggregate vector 175. The estimation error is, for example, the L1 norm error, which represents the absolute value of the difference between two values.

[0102] The information processing device 100 sorts the multiple dimensions of the aggregate vector 175 in descending order of estimation error. This estimation error is standardized based on the variance of the aggregate vector 172 used in the training phase. The information processing device 100 selects a certain number of top dimensions with large estimation errors and generates an effective dimension list 176 indicating the selected dimensions. For example, the information processing device 100 selects 100 dimensions. The effective dimension list 176 is, for example, an index list enumerating the indices of the selected dimensions. However, the effective dimension list 176 may also enumerate the indices of dimensions that were not selected.

[0103] During the operational phase, the information processing device 100 acquires an aggregation vector 177. The aggregation vector 177 is calculated from the time-series data of the input period using the method described above. The information processing device 100 converts the aggregation vector 177 into a filtered aggregation vector 178 using the effective dimension list 176. The filtered aggregation vector 178 includes the effective dimensions indicated in the effective dimension list 176 among the multiple dimensions included in the aggregation vector 177, and does not include other dimensions. By filling in the space of the deleted dimensions, the number of dimensions of the filtered aggregation vector 178 becomes smaller than that of the aggregation vector 177.

[0104] The information processing device 100 converts the filtered aggregate vector 178 into an additional feature vector 179 by performing dimensionality reduction on the filtered aggregate vector 178. The additional feature vector 179 is a vector with fewer dimensions than the filtered aggregate vector 178. For example, the number of dimensions of the additional feature vector 179 is approximately 1 to 10. Dimensionality reduction may be performed using principal component analysis or an autoencoder.

[0105] During machine learning to train the accuracy estimation model 132, the information processing device 100 calculates a set of filtered aggregate vectors corresponding to the sample set. The information processing device 100 can then calculate a projection matrix for transforming each filtered aggregate vector through principal component analysis. By saving the projection matrix calculated during machine learning, the information processing device 100 can transform the filtered aggregate vectors when using the accuracy estimation model 132. However, the information processing device 100 may obtain a general-purpose projection matrix corresponding to the pre-compression and post-compression dimensions from another information processing device.

[0106] Furthermore, the information processing device 100 may train an autoencoder using a set of filtered aggregate vectors corresponding to the sample set. An autoencoder is a neural network that converts an input vector into a latent space vector with fewer dimensions than the input vector. By storing the autoencoders generated during machine learning, the information processing device 100 can convert the filtered aggregate vectors when using the accuracy estimation model 132. However, the information processing device 100 may obtain a general-purpose autoencoder corresponding to the number of dimensions before and after compression from another information processing device.

[0107] The information processing device 100 determines the dimensions with large estimation errors during machine learning to train the accuracy estimation model 132. Therefore, as a rule, after machine learning, the information processing device 100 extracts the determined fixed dimensions from the aggregation vector 177. However, when using the accuracy estimation model 132, the user may tune the related estimation model 134 using the user's own time-series data. The user may also update the effective dimension list 176 using the tuned related estimation model 134. Next, the functions of the information processing device 100 and the processing procedures of the information processing device 100 will be described.

[0108] Figure 10 is a block diagram showing an example of the functions of an information processing device. The information processing device 100 includes a data storage unit 121, a model storage unit 122, a feature generation unit 123, an accuracy estimation unit 124, a time series prediction unit 125, a base model training unit 126, a related model training unit 127, and an accuracy model training unit 128.

[0109] The data storage unit 121 and the model storage unit 122 are implemented using, for example, RAM 102 or HDD 103. The feature generation unit 123, accuracy estimation unit 124, time series prediction unit 125, base model training unit 126, related model training unit 127, and accuracy model training unit 128 are implemented using, for example, CPU 101, GPU 104, and program.

[0110] The data storage unit 121 stores a sample set of time-series data. The sample set includes a large number of samples representing various types of time-series data. The information processing device 100 may increase the number of samples using data augmentation techniques.

[0111] The model storage unit 122 stores trained machine learning models. The model storage unit 122 stores multiple prediction models, accuracy estimation models 132, base models 133, and related estimation models 134. The multiple prediction models may include statistical models, deep machine learning models, or machine learning models other than deep machine learning models.

[0112] The feature generation unit 123 receives time-series data and calculates the features of the received time-series data. The feature generation unit 123 calculates statistical features from the time-series data using a specific algorithm. The feature generation unit 123 also inputs the time-series data into a base model, extracts intermediate vectors from the base model, and calculates aggregate vectors. The feature generation unit 123 also transforms the aggregate vectors to generate additional features. Finally, the feature generation unit 123 concatenates the statistical features and additional features to generate input features.

[0113] The accuracy estimation unit 124 receives input features. The accuracy estimation unit 124 inputs the received input features into the accuracy estimation model 132 and obtains estimated prediction accuracy values ​​for each of the multiple prediction models from the accuracy estimation model 132.

[0114] The time series forecasting unit 125 accepts time series data for the input period and a specification of the forecasting model to be used. The time series forecasting unit 125 inputs the accepted time series data into the specified forecasting model and obtains time series data for the forecast period from the forecasting model. The time series forecasting unit 125 may also accept the forecast accuracy of multiple forecasting models. In that case, the time series forecasting unit 125 selects some of the forecasting models according to their forecast accuracy. The time series forecasting unit 125 inputs the accepted time series data into each of the selected forecasting models, obtains time series data for the forecast period from each forecasting model, and synthesizes the obtained time series data to generate the final forecast result.

[0115] The base model training unit 126 trains the base model 133 using machine learning. The base model training unit 126 inputs time series data samples into the base model 133, evaluates the error of the output of the base model 133, and updates the parameter values ​​using backpropagation. The base model training unit 126 may input the entire sample into the base model 133. Alternatively, the base model training unit 126 may divide the sample into an input period and a prediction period, and input only the input period into the base model 133.

[0116] The association model training unit 127 trains the association estimation model 134 using machine learning. The association model training unit 127 obtains the statistical features and aggregate vector of the sample. The association model training unit 127 determines the parameter values ​​of the association estimation model 134 so that the association estimation model 134 estimates the values ​​of each dimension of the aggregate vector from the statistical features. The association model training unit 127 evaluates the estimation error of each dimension of the trained association estimation model 134 using the sample. The association model training unit 127 generates an effective dimension list indicating the dimensions with large estimation errors. The effective dimension list is stored in the model storage unit 122.

[0117] The related model training unit 127 may use the same sample set as the base model training unit 126, or it may use a different sample set than the base model training unit 126. Furthermore, the related model training unit 127 may use statistical features calculated from the entire sample, or it may use aggregate vectors calculated from the entire sample. Also, the related model training unit 127 may divide the sample into an input period and a prediction period, and use statistical features calculated only from the input period, or it may use aggregate vectors calculated only from the input period.

[0118] The accuracy model training unit 128 trains the accuracy estimation model 132 using machine learning. The accuracy model training unit 128 divides the sample into an input period and a prediction period. The accuracy model training unit 128 obtains input features generated from the time series data of the input period. The accuracy model training unit 128 also obtains the prediction results of each of the multiple prediction models, compares the prediction results with the ground truth, and calculates the prediction accuracy of each of the multiple prediction models.

[0119] The accuracy model training unit 128 inputs the input features into the accuracy estimation model 132 and obtains an estimated value of the prediction accuracy from the accuracy estimation model 132. The accuracy model training unit 128 compares the estimated value with the measured value and calculates the estimation error for each of the multiple prediction models. The accuracy model training unit 128 updates the parameter values ​​of the accuracy estimation model 132 so that the estimation error is reduced. The accuracy model training unit 128 may use the same sample set as the base model training unit 126, or it may use a different sample set than the base model training unit 126. Also, the accuracy model training unit 128 may use the same sample set as the related model training unit 127, or it may use a different sample set than the related model training unit 127.

[0120] The base model training unit 126 may display the trained base model 133 on the display device 111, or transmit the trained base model 133 to another information processing device. The related model training unit 127 may display the trained related estimation model 134 on the display device 111, or transmit the trained related estimation model 134 to another information processing device. The related model training unit 127 may display the effective dimension list on the display device 111, or transmit the effective dimension list to another information processing device.

[0121] Furthermore, the accuracy model training unit 128 may display the trained accuracy estimation model on the display device 111, or transmit the trained accuracy estimation model to another information processing device. Furthermore, the accuracy estimation unit 124 may display the estimated prediction accuracy on the display device 111, or transmit the estimated prediction accuracy to another information processing device. Furthermore, the time series prediction unit 125 may display the predicted time series data on the display device 111, or transmit the predicted time series data to another information processing device.

[0122] Figure 11 is a flowchart illustrating an example of a machine learning procedure. In step S10, the base model training unit 126 trains the base model 133 using a sample set of time-series data. If a pre-trained base model 133 exists, the base model training unit 126 may omit step S10. In step S11, the related model training unit 127 selects one unselected sample from the sample set. The feature generation unit 123 calculates the statistical features of the selected sample. The sample set used in step S11 may be different from the sample set used in step S10.

[0123] In step S12, the feature generation unit 123 uses the base model 133 trained in step S10 to convert the selected samples into a set of intermediate vectors. Here, the feature generation unit 123 inputs the samples into the base model 133 and extracts from the base model 133 multiple intermediate vectors corresponding to multiple time values ​​contained in the samples.

[0124] In step S13, the feature generation unit 123 converts the set of intermediate vectors into an aggregate vector. Here, the feature generation unit 123 calculates one or more statistics for each dimension of the intermediate vectors and generates an aggregate vector that enumerates the statistics for multiple dimensions. In step S14, the related model training unit 127 adds the pairs of statistical features from step S11 and the aggregate vectors from step S13 to the training data. In step S15, the related model training unit 127 determines whether all samples have been selected. If all samples have been selected, the process proceeds to step S16. If there are unselected samples, the process returns to step S11.

[0125] In step S16, the related model training unit 127 divides the training data into N subsets. For example, N = 5. Preferably, the N subsets contain an equal number of records. Each record contains one statistical feature and one aggregation vector. In steps S16 to S21, the related model training unit 127 performs cross-validation. In step S17, the related model training unit 127 selects one test subset from the N subsets and combines the remaining subsets.

[0126] In step S18, the association model training unit 127 trains the association estimation model 134 using a combined subset. For example, the association model training unit 127 generates multiple multivariate regression models corresponding to multiple dimensions of the aggregate vector. Each multivariate regression model estimates the value of one dimension of the aggregate vector from statistical features.

[0127] In step S19, the related model training unit 127 calculates the estimated error for each dimension using the test subset selected in step S17. In step S20, the related model training unit 127 determines whether all subsets have been selected as test subsets. If all subsets have been selected, the process proceeds to step S21. If there are unselected subsets, the process returns to step S17. In step S21, the related model training unit 127 selects the median of N estimated errors for each dimension. From this point onward, the median selected here is used as the estimated error for that dimension.

[0128] Figure 12 is a flowchart (continued 1) showing an example of the machine learning procedure. In step S22, the related model training unit 127 sorts the dimensions of the aggregate vector in descending order of estimation error. In step S23, the related model training unit 127 generates an effective dimension list that shows a certain number of dimensions with large estimation errors.

[0129] In step S24, the accuracy model training unit 128 selects one unselected sample from the sample set and divides the selected sample into an input period and a prediction period. The length of the input period or the prediction period may be specified by the user. In step S25, the feature generation unit 123 calculates the statistical features of the input period of the sample.

[0130] In step S26, the feature generation unit 123 uses the base model 133 to convert the input period of the samples into a set of intermediate vectors. In step S27, the feature generation unit 123 converts the set of intermediate vectors into an aggregate vector. In step S28, the feature generation unit 123 uses the effective dimension list from step S23 to remove some dimensions from the aggregate vector. As a result, the feature generation unit 123 converts the aggregate vector into a filtered aggregate vector with fewer dimensions than the aggregate vector.

[0131] In step S29, the accuracy model training unit 128 determines whether all samples have been selected. If all samples have been selected, the process proceeds to step S30. If there are unselected samples, the process returns to step S24. In step S30, the feature generation unit 123 converts each filtered aggregate vector into an additional feature by dimensionality reduction.

[0132] For example, the feature generation unit 123 performs principal component analysis on the set of filtered aggregate vectors corresponding to the sample set and calculates projection matrices corresponding to the pre-compression and post-compression dimensions. The feature generation unit 123 calculates additional features by multiplying each filtered aggregate vector by the calculated projection matrices. The feature generation unit 123 also stores the calculated projection matrices. In step S31, the feature generation unit 123 generates input features by concatenating the statistical features from step S25 and the additional features from step S30.

[0133] Figure 13 is a flowchart (continued 2) showing an example of the machine learning procedure. In step S32, the accuracy model training unit 128 selects one unselected sample from the sample set. The time series forecasting unit 125 inputs the input period of the sample into each forecasting model and predicts the forecast period of the sample. In step S33, the accuracy model training unit 128 compares the prediction with the correct answer and calculates the prediction accuracy of each forecasting model.

[0134] In step S34, the accuracy model training unit 128 adds pairs of input features and prediction accuracy sets to the training data for the selected samples. The prediction accuracy set is a collection of prediction accuracies from multiple prediction models. In step S35, the accuracy model training unit 128 determines whether all samples have been selected. If all samples have been selected, the process proceeds to step S36. If there are unselected samples, the process returns to step S32.

[0135] In step S36, the accuracy model training unit 128 trains the accuracy estimation model 132 using the training data. Here, the accuracy model training unit 128 inputs the input features into the accuracy estimation model 132 and obtains the estimated value of the predicted accuracy set from the accuracy estimation model 132. The accuracy model training unit 128 evaluates the error between the estimated value of the predicted accuracy set and the measured value, and updates the parameter values ​​of the accuracy estimation model 132 to reduce the error. In step S37, the accuracy model training unit 128 outputs the accuracy estimation model 132.

[0136] Figure 14 is a flowchart illustrating an example of the time series forecasting procedure. In step S40, the time series forecasting unit 125 receives time series data for the input period. In step S41, the feature generation unit 123 calculates statistical features of the time series data for the input period. In step S42, the feature generation unit 123 uses the base model 133 to convert the time series data for the input period into an intermediate vector set.

[0137] In step S43, the feature generation unit 123 converts the set of intermediate vectors into an aggregate vector. In step S44, the feature generation unit 123 removes some dimensions from the aggregate vector using the effective dimension list. This converts the aggregate vector into a filtered aggregate vector. In step S45, the feature generation unit 123 converts the filtered aggregate vector into additional features by dimensionality reduction. For example, the feature generation unit 123 calculates additional features by multiplying the filtered aggregate vector by the projection matrix calculated in machine learning.

[0138] In step S46, the feature generation unit 123 generates input features by concatenating the statistical features from step S41 and the additional features from step S45. In step S47, the accuracy estimation unit 124 estimates the prediction accuracy of multiple prediction models from the input features using the accuracy estimation model 132. In step S48, the time series prediction unit 125 predicts the time series data for the prediction period from the time series data for the input period using the prediction model with the highest prediction accuracy. In step S49, the time series prediction unit 125 outputs the time series data for the prediction period. Next, the relationship between the design of the input features and the accuracy of the accuracy estimation model 132 will be explained.

[0139] Figure 15 is a graph showing a comparative example of accuracy estimation. In this comparative example, the information processing device 100 employs a Light Gradient Boosting Machine as the model structure for the accuracy estimation model 132. The information processing device 100 also uses Seasonal Naive, Auto ETS, Patch TST, Temporal Fusion Transformer, Deep AR, DLinear, TiDE, WaveNet, and Simple Feed Forward as prediction models. However, the information processing device 100 uses Seasonal Naive as the reference prediction model for accuracy comparison. Therefore, the information processing device 100 evaluates the accuracy estimation for the eight prediction models.

[0140] Furthermore, the prediction accuracy estimated by the accuracy estimation model 132 is MASE. Also, the aggregate vector contains only the mean, and the number of dimensions of the aggregate vector is 768. The information processing device 100 extracts 100 dimensions from the 768 dimensions of the aggregate vector in which the estimation error of the association estimation model 134 is large. The association estimation model 134 is a collection of multivariate regression models. The estimation error of the association estimation model 134 is the L1 norm.

[0141] Dimensionality reduction is performed using principal component analysis. The information processing device 100 tries 10 different dimensionalities for the additional features, ranging from 1 to 10 dimensions. The information processing device 100 also generates an accuracy estimation model 132 from each of the six datasets. Therefore, for each of the 10 different dimensionalities, the information processing device 100 evaluates the accuracy of 48 different accuracy estimations (8 prediction models × 6 datasets).

[0142] The indicator of accuracy is the Area Under the Curve (AUC) of the Receiver Operating Characteristic (ROC) curve. The information processing device 100 compares the estimation results of the accuracy estimation model 132 with the correct answer to determine whether the prediction accuracy of each of the eight prediction models is higher than that of the reference prediction model. This ROC-AUC indicates how close the estimation result is to the correct answer.

[0143] Graph 164 shows the relationship between the design of the input features and the accuracy of the accuracy estimation model 132. In Graph 164, S represents the case where the input features do not include additional features. L represents the case where additional features are calculated using the values ​​of the dimension with the smallest estimation error in the aggregate vector. H represents the case where additional features are calculated using the values ​​of the dimension with the largeest estimation error in the aggregate vector, and corresponds to the method of the second embodiment. T represents the case where additional features are calculated using the values ​​of all dimensions included in the aggregate vector.

[0144] The bars labeled S represent the percentage of evaluation results that had the highest accuracy (S) out of 48 possible evaluation results. The bars labeled L represent the percentage of evaluation results that had the highest accuracy (L) out of 48 possible evaluation results. The bars labeled H represent the percentage of evaluation results that had the highest accuracy (H) out of 48 possible evaluation results. The bars labeled T represent the percentage of evaluation results that had the highest accuracy (T) out of 48 possible evaluation results. The sum of the percentages for S, L, H, and T is 100%.

[0145] As shown in Graph 164, using additional features significantly improves the accuracy of accuracy estimation compared to using only statistical features. Furthermore, extracting dimensions with large estimation errors in the association estimation model 134 significantly improves the accuracy of accuracy estimation compared to extracting dimensions with small estimation errors. Additionally, extracting dimensions with large estimation errors improves the accuracy of accuracy estimation compared to using all dimensions of the aggregate vector.

[0146] In particular, except when the number of dimensions of the additional features is 1, the smaller the number of dimensions, the greater the difference between the two. Therefore, the information processing device 100 can achieve high accuracy even with limited computing resources by extracting dimensions with large estimation errors from the aggregate vector. As a result, the information processing device 100 can quickly estimate the prediction accuracy of the prediction model.

[0147] As described above, the information processing device 100 of the second embodiment predicts time series data for a period beyond the input time series data using multiple prediction models. This allows the information processing device 100 to use a prediction model suitable for the properties of the input time series data, thereby improving prediction accuracy compared to using a single prediction model. Furthermore, the information processing device 100 calculates the feature quantities of the input time series data and estimates the prediction accuracy of each prediction model using the accuracy estimation model 132 before executing each prediction model. This allows the information processing device 100 to omit the execution of unnecessary prediction models that are likely to have low prediction accuracy, thereby speeding up time series forecasting.

[0148] Furthermore, in addition to statistical features, the information processing device 100 calculates additional features from time-series data using the base model 133. The information processing device 100 generates input features by concatenating the statistical features and the additional features, and inputs these input features into the accuracy estimation model 132. The additional features based on the base model 133 have high compatibility with deep machine learning type prediction models. Therefore, the information processing device 100 can estimate the prediction accuracy of various types of prediction models, including statistical models and deep machine learning models, with high accuracy.

[0149] Furthermore, the information processing device 100 generates an association estimation model 134 that has learned the relationship between statistical features and multiple dimensions included in the aggregation vector. The information processing device 100 uses the association estimation model 134 to determine dimensions with low similarity to the statistical features. The information processing device 100 extracts the values ​​of the dimensions with low similarity and generates additional features. As a result, the information processing device 100 can efficiently improve accuracy even with small additional features.

[0150] The above merely illustrates the principle of the present invention. Furthermore, numerous modifications and changes are possible for those skilled in the art, and the present invention is not limited to the exact configurations and applications shown and described above. All corresponding modifications and equivalents are considered to be within the scope of the present invention as defined by the appended claims and their equivalents.

[0151] 10 Information processing device 11 Memory unit 12 Processing unit 13 Prediction model 14 Machine learning model 15 Accuracy estimation model 16 Time series data 17a, 17a Features 18 Input features 19 Prediction accuracy

Claims

1. A method for generating an accuracy estimation model, in which a computer performs the following steps: calculate a first feature from time series data using a specific algorithm; input the time series data into a trained machine learning model to obtain a second feature calculated from the time series data using the trained machine learning model; input the time series data into a prediction model to evaluate the prediction accuracy of data for a period later than the time series data, as predicted by the prediction model; and generate an accuracy estimation model that learns the relationship between input features based on the first and second features and the prediction accuracy.

2. The method for generating an accuracy estimation model according to claim 1, wherein the trained machine learning model is a neural network comprising multiple layers, and the acquisition process includes extracting vector data from an intermediate layer among the multiple layers and generating the second feature from the vector data.

3. The method for generating an accuracy estimation model according to claim 1, wherein the second feature is vector data including multiple dimensions, and the computer further performs a process to convert the second feature into a third feature having fewer dimensions than the second feature by deleting the values ​​of some of the multiple dimensions from the second feature, and the input feature includes the first feature and the third feature.

4. The method for generating an accuracy estimation model according to claim 3, wherein the computer further performs a process of generating an association estimation model that learns the relationship between the first feature and the multiple dimensions included in the second feature, and determining some of the dimensions using the association estimation model.

5. The method for generating an accuracy estimation model according to claim 4, wherein the process of determining includes evaluating the estimation accuracy of the associated estimation model for each of the multiple dimensions, and preferentially selecting the dimensions with high estimation accuracy from among the multiple dimensions as some of the dimensions.

6. A program for generating an accuracy estimation model that causes a computer to perform the following steps: calculate a first feature from time series data using a specific algorithm; input the time series data into a trained machine learning model to obtain a second feature calculated from the time series data using the trained machine learning model; input the time series data into a prediction model to evaluate the prediction accuracy of data for a period later than the time series data, as predicted by the prediction model; and generate an accuracy estimation model that learns the relationship between input features based on the first and second features and the prediction accuracy.