Information processing device and information processing method

WO2026203243A1PCT designated stage Publication Date: 2026-10-01NT T INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/012620
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2026-10-01

Smart Images

  • Figure JP2025012620_01102026_PF_FP_ABST
    Figure JP2025012620_01102026_PF_FP_ABST
Patent Text Reader

Abstract

An information processing device 10 includes a control unit 11 that acquires, as time series data, a first feature indicating a feature on a first temporal scale and a second feature indicating a feature on a second temporal scale, matches the temporal scales of the first feature and the second feature, inputs, to a trained model, the first feature and the second feature after the temporal scales are matched and the first feature and the second feature before the temporal scales are matched, and predicts future temporal variations of the time series data.
Need to check novelty before this filing date? Find Prior Art

Description

Information Processing Apparatus and Information Processing Method

[0001] The present disclosure relates to an information processing apparatus and an information processing method.

[0002] In time series forecasting that predicts future time series data from past time series data, long-term forecasting is an important issue. For example, in time series forecasting such as traffic forecasting and energy demand forecasting, it is necessary to accurately capture not only short-term fluctuations but also long-term trends.

[0003] Non-Patent Document 1 describes a technology called DRFormer (Dynamic Receptive Field Transformer) that captures short-term features and long-term features. DRFormer performs an operation of aligning the data size of each scale through Hierarchical Max Pooling. is described.

[0004] Ding, Ruixin, et al. "DRFormer: Multi-Scale Transformer Utilizing Diverse Receptive Fields for Long Time-Series Forecasting." Proceedings of the 33rd ACM International Conference on Information and Knowledge Management. 2024.

[0005] However, in the above-mentioned DRFormer, when features of time series data for each scale are compressed in Hierarchical Max Pooling, information loss may occur in the time series data of each scale. Such information loss causes a limit to long-term prediction accuracy.

[0006] As described above, the conventional configuration related to time series forecasting cannot accurately capture both short-term fluctuations and long-term trends, and there is room for improvement in the accuracy of long-term forecasting.

[0007] The purpose of this disclosure is to enable more accurate long-term forecasting of time series data.

[0008] According to this disclosure, the information processing device includes a control unit that acquires a first feature representing a feature on a first temporal scale and a second feature representing a feature on a second temporal scale of time series data that shows values ​​according to the passage of time, aligns the temporal scales of the first feature and the second feature, and inputs the first feature and the second feature after the temporal scale alignment and the first feature and the second feature before the temporal scale alignment into a trained model to predict future temporal fluctuations of the time series data.

[0009] According to one embodiment of this disclosure, long-term forecasting of time series data can be performed with higher accuracy.

[0010] This is a block diagram showing an example of the hardware configuration of an information processing device according to one embodiment. This is a block diagram showing an example of the functional configuration of an information processing device according to one embodiment. This is a diagram showing an example of time-series data. This is a diagram showing an example of time-series data organized by day. This is a diagram showing an example of time-series data organized by week. This is a diagram showing an example of time-series data organized by month. This is a flowchart showing an example of the operation of an information processing device.

[0011] Hereinafter, an embodiment of the present disclosure will be described with reference to the drawings. In each drawing, parts having the same configuration or function are denoted by the same reference numerals. In the description of this embodiment, redundant descriptions of the same parts may be omitted or simplified as appropriate.

[0012] <Hardware Configuration of Information Processing Device 10> Figure 1 is a block diagram showing an example of the hardware configuration of an information processing device 10 according to one embodiment. The information processing device 10 predicts future temporal fluctuations of time-series data based on time-series data represented by multiple temporal scales. In the following explanation, the time-series data is described as data showing temporal fluctuations in the amount of communication traffic (Mbps) in a certain network, but the content of the data shown by the time-series data is arbitrary. For example, the time-series data may be data showing temporal fluctuations in energy demand (e.g., electricity).

[0013] The information processing device 10 is one or a plurality of computer devices capable of communicating with each other. Specifically, the information processing device 10 may be any general-purpose electronic device such as a WS (Workstation), PC (Personal Computer), or tablet terminal, or it may be another dedicated electronic device. As shown in Figure 1, the information processing device 10 comprises a control unit 11, a storage unit 12, a communication unit 13, an input unit 14, and an output unit 15.

[0014] The control unit 11 includes one or more processors. In one embodiment, the "processor" is a general-purpose processor or a dedicated processor specialized for a specific process, but is not limited to these. The control unit 11 is communicatively connected to each component constituting the information processing device 10 and controls the operation of the entire information processing device 10.

[0015] The storage unit 12 includes, for example, any storage module such as an HDD (Hard Disk Drive), SSD (Solid State Drive), ROM (Read-Only Memory), and RAM (Random Access Memory). The storage unit 12 may function as, for example, main memory, auxiliary memory, or cache memory. The storage unit 12 stores any information used in the operation of the information processing device 10. For example, the storage unit 12 may store system programs, application programs, and various information received by the communication unit 13. The storage unit 12 is not limited to those built into the information processing device 10, but may also be an external database or an external storage module.

[0016] The memory unit 12 may store time-series data that shows values ​​according to the passage of time. The memory unit 12 may store the same time-series data expressed on multiple time scales, such as daily, weekly, and monthly. The memory unit 12 may store a trained model for predicting future temporal fluctuations of time-series data based on the time-series data expressed on multiple time scales. The trained model may be a model that has been trained in advance using machine learning techniques, such as the parameters of a neural network.

[0017] The communication unit 13 includes an optional communication module that can communicate with other devices such as network storage using any communication technology. The communication unit 13 may further include a communication control module for controlling communication with other devices, and a storage module for storing communication data such as identification information necessary for communication with other devices.

[0018] The input unit 14 includes one or more input interfaces that receive user input operations and acquire input information based on user operations. For example, the input unit 14 may be, but is not limited to, physical keys, capacitive keys, a pointing device, a touchscreen integrated with the display of the output unit 15, or a microphone that accepts voice input.

[0019] The output unit 15 includes one or more output interfaces that output information to the user and notify the user. For example, the output unit 15 is a monitor (display) that outputs information as an image, or a speaker that outputs information as sound, but is not limited to these. Such a monitor may be, for example, a liquid crystal panel display or an organic EL (Electro Luminescence) display. At least one of the above-mentioned input unit 14 and output unit 15 may be configured integrally with the information processing device 10, or may be provided as a separate unit.

[0020] The information processing device 10 can also be realized by a computer and a program, and the program can be recorded on a recording medium or provided via a network. In other words, the functions of the information processing device 10 can be realized by executing a computer program (program) according to this embodiment on the processor included in the control unit 11. To put it another way, the functions of the information processing device 10 can be realized by software. The computer program causes the computer to execute the processing steps included in the operation of the information processing device 10, thereby realizing the functions corresponding to the processing of each step on the computer. In other words, the computer program is a program that causes the computer to function as the information processing device 10 according to this embodiment.

[0021] Computer programs can be recorded on computer-readable recording media. Examples of computer-readable recording media include magnetic recording devices, optical discs, magneto-optical recording media, or semiconductor memory. Programs can be distributed, for example, by selling, transferring, or leasing portable recording media such as DVDs (Digital Versatile Discs) or USB (Universal Serial Bus) memory containing the programs. Programs may also be distributed by storing them in server storage and transferring them from the server to other computers via a network. Programs may also be provided as program products.

[0022] A computer may, for example, temporarily store a program recorded on a portable storage medium or a program transferred from a server in its main memory. The computer may then read the program stored in the main memory with its processor and execute the processing according to the read program. The computer may also directly read a program from a portable storage medium and execute the processing according to the program. The computer may sequentially execute the processing according to the received program each time a program is transferred to it from a server. Such processing may also be performed by a so-called ASP-type service that does not transfer programs from the server to the computer, but realizes its function only through execution instructions and result acquisition. "ASP" is an abbreviation for Application Service Provider. A program includes information used for processing by an electronic computer that is equivalent to a program. For example, data that is not a direct instruction to the computer but has the nature of defining the computer's processing falls under "equivalent to a program".

[0023] Some or all of the functions of the information processing device 10 may be implemented by a dedicated circuit included in the control unit 11. In other words, some or all of the functions of the information processing device 10 may be implemented by hardware. Furthermore, the information processing device 10 may be implemented by a single computer or by the cooperation of multiple computers. For example, the information processing device 10 may be implemented by a server device installed in, for example, a data center or cloud computing system, and a terminal device that accesses the server device via a network.

[0024] In the configuration described above, the information processing device 10 prevents information loss between time scales and enables learning of short-term and long-term features by uniformly modeling the features of time series data represented by different time scales (e.g., daily, weekly, monthly). Specifically, the information processing device 10 acquires a first feature that shows the features of the time series data, which shows values ​​according to the passage of time, at a first time scale (e.g., daily), and a second feature that shows the features at a second time scale (e.g., monthly). The information processing device 10 aligns the time scales of the first and second features. The information processing device 10 inputs the first and second features after the time scales have been aligned, and the first and second features before the time scales have been aligned, into a trained model to predict future temporal fluctuations of the time series data.

[0025] In this way, the information processing device 10 integrates the features of multiple time series data with different temporal scales after temporal scale alignment using the residual connection method, and then trains a model. The information processing device 10 handles features of different temporal scales (daily, weekly, monthly) in a unified manner, prevents information loss between temporal scales, and handles short-term and long-term features complementaryly. Therefore, the information processing device 10 makes it possible to accurately capture both short-term fluctuations and long-term trends, and to perform long-term forecasts of time series data with higher accuracy.

[0026] <Functional Configuration of Information Processing Device 10> Figure 2 is a block diagram showing an example of the functional configuration of an information processing device 10 according to one embodiment. As shown in Figure 2, the information processing device 10 comprises the functional elements of a feature extraction unit 21, a scale matching unit 22, and a learning unit 23. The learning unit 23 comprises an intermediate processing layer 231 and a post-processing layer 233. At least some of these functional elements may be implemented by software or by hardware.

[0027] The feature extraction unit 21 extracts features from time-series data across multiple time scales. For example, the feature extraction unit 21 may receive daily data 321, weekly data 322, and monthly data 323 obtained from the same time-series data 31, and extract features (e.g., daily features, weekly features, and monthly features) from each of these data sets. The feature extraction unit 21 outputs the extracted features for each time scale to the scale matching unit 22.

[0028] The scale matching unit 22 aligns the temporal scales of multiple features input from the feature extraction unit 21. For example, the scale matching unit 22 may upsample or downsample the time axis to convert the features obtained by the feature extraction unit 21 for each temporal scale into a size that can be integrated, thereby converting them into a common feature dimension. The scale matching unit 22 outputs the multiple features with aligned temporal scales to the learning unit 23.

[0029] The learning unit 23 outputs a prediction result 35 that shows the future temporal fluctuations of the time series data 31, based on a plurality of features input from the scale matching unit 22. For example, the learning unit 23 may include an intermediate processing layer 231 and a post-processing layer 233. As shown in Figure 2, the learning unit 23 includes an adder 232 that inputs the output of the feature extraction unit 21 to the post-processing layer 233, skipping the scale matching unit 22 and the intermediate processing layer 231. In this way, the learning unit 23 outputs the prediction result 35 using not only a plurality of features after the temporal scale has been matched in the scale matching unit 22, but also a plurality of features before the temporal scale has been matched. That is, the learning unit 23 implements a learning model that has the function of preventing information loss between temporal scales by complementaryly learning the short-term fluctuations and long-term trends of the time series data 31 based on the features matched in the scale matching unit 22. The learning unit 23 may also output not only the prediction results, but also the characteristics of the time series data 31 for each time scale (for example, daily data 321, weekly data 322, monthly data 323).

[0030] In this embodiment, as an example, the information processing device 10 processes time-series data 31 organized on three time scales: daily, weekly, and monthly. However, the types of time scales are not limited to these. The information processing device 10 is applicable to any combination and number of time scales.

[0031] Next, we will explain the data in Figure 2 and the details of each functional element.

[0032] (Time-series data 31) Figure 2 shows an example in which daily data 321, weekly data 322, and monthly data 323 obtained from time-series data 31 are input to the feature extraction unit 21. The information processing device 10 may statically divide the pre-recorded time-series data 31 into daily, weekly, and monthly time scales to obtain daily data 321, weekly data 322, and monthly data 323. Alternatively, the information processing device 10 may obtain daily data 321, weekly data 322, and monthly data 323 that have been divided and obtained in advance by another device.

[0033] Figure 3 shows an example of time-series data 31. In the example in Figure 3, the time-series data 31 shows the hourly traffic volume (Mbps) in a certain network. Figure 4 shows an example of time-series data organized by day (daily data 321). Figure 5 shows an example of time-series data organized by week (weekly data 322). Figure 6 shows an example of time-series data organized by month (monthly data 323).

[0034] The information processing device 10 may calculate a daily average value for time-series data 31 recorded as hourly communication traffic volume to obtain daily communication traffic volume and acquire daily data 321. For example, the information processing device 10 may calculate the data for January 1st as the average value of the traffic volume from 0:00 to 23:00 on January 1st. The information processing device 10 may perform such processing on a daily basis to acquire daily data 321.

[0035] The information processing device 10 may calculate a weekly average value for the daily data 321, which shows the amount of communication traffic each day, to obtain the weekly amount of communication traffic and acquire weekly data 322. For example, the information processing device 10 may calculate the data for the first week as the average traffic amount for each day from January 1st to January 7th. The information processing device 10 may perform such processing on a weekly basis to acquire weekly data 322.

[0036] The information processing device 10 may calculate a monthly average value for weekly data 322 showing weekly communication traffic volume to obtain monthly communication traffic volume and acquire monthly data 323. For example, the information processing device 10 may calculate the data for January as the average traffic volume from the first week to the fourth week of January. The information processing device 10 may perform such processing on a monthly basis to acquire monthly data 323.

[0037] (Feature Extraction Unit 21) The feature extraction unit 21 extracts features from the daily data 321, weekly data 322, and monthly data 323, respectively. The features may be statistical data such as mean, maximum, minimum, and variance. The feature extraction unit 21 may output features on the same temporal scale as the input time series data (daily data 321, weekly data 322, and monthly data 323). The feature extraction unit 21 may be implemented as a pre-trained model that has been trained using machine learning methods such as neural networks. For example, the feature extraction unit 21 may be implemented using Mamba blocks, a Self-attention layer of a Transformer, a convolutional layer, etc.

[0038] In this embodiment, the feature extraction unit 21 receives time-series data divided into multiple time scales (for example, daily data 321, weekly data 322, and monthly data 323) and extracts features using a common configuration for each time scale, but is not limited to this configuration. For example, the feature extraction unit 21 may have a configuration for receiving time-series data 31 that is not divided into multiple time scales and outputting features for each of the multiple time scales.

[0039] For example, the feature extraction unit 21 may extract daily features from the time series data 31 using a pre-trained convolutional neural network (CNN). By using a convolutional neural network, the feature extraction unit 21 can capture rapid fluctuations and periodicity in the time series data 31 over short periods. Specifically, for example, the feature extraction unit 21 can extract local patterns in the time series data 31 by using a one-dimensional convolutional layer (1D-CNN) and obtain useful features from the hourly traffic volume for each day. It should be noted that the feature extraction unit 21 can also extract short-term trends in the time series data 31 as features by using, for example, a recurrent neural network (RNN) and long-term short-term memory (LSTM), not just a convolutional neural network.

[0040] Furthermore, for example, the feature extraction unit 21 may extract weekly features from the time series data 31 by applying a Mamba block that uses a pre-trained linear state-space model (SSM). By using the Mamba block, the feature extraction unit 21 can capture the periodicity of the time series data 31 on a day-of-the-week basis and the fluctuations on a weekly basis. Specifically, for example, the feature extraction unit 21 can obtain longer-term dependencies from the time series data 31 using the Mamba block and extract weekly fluctuation patterns while suppressing the effects of short-term noise. Note that the feature extraction unit 21 can also extract weekly periodicity as a feature if, for example, the encoder layer and convolutional layer of a Transformer are used, not limited to the Mamba block.

[0041] Furthermore, for example, the feature extraction unit 21 may extract monthly features from the time series data 31 by using a pre-trained self-attention mechanism. By using a self-attention mechanism, the feature extraction unit 21 can extract features that reflect longer-term trends and seasonal fluctuations in the time series data 31. Specifically, for example, the feature extraction unit 21 may apply an encoder block of a self-attention mechanism such as a Transformer to extract features that take into account all past data. It should be noted that the feature extraction unit 21 can also extract long-term dependencies as monthly features by using, for example, a recursive network, a linear state-space model, or a convolutional network, not limited to a self-attention mechanism.

[0042] (Scale matching unit 22) The scale matching unit 22 unifies the dimensions of the time scale by matching the time axis of the features of the time series data 31 at each time scale through upsampling or downsampling. The scale matching unit 22 may be implemented as a pre-trained model that has been trained in advance using machine learning methods such as neural networks.

[0043] For example, on a daily temporal scale, the temporal resolution of data is high, so it is necessary to reduce the temporal resolution (downsampling) when integrating with a weekly or monthly temporal scale. For example, the scale matching unit 22 can reduce the temporal resolution by applying a convolutional layer or average pooling to aggregate daily features belonging to a certain temporal range. Further, the scale matching unit 22 can also adjust the daily data 321 to the resolution of a weekly or monthly temporal scale by performing linear interpolation in the time axis direction, for example.

[0044] Further, since the weekly temporal scale has an intermediate temporal resolution between daily data and monthly data, conversion adaptable to both daily and monthly scales is performed during integration. For example, when integrating weekly features into the daily scale, the scale matching unit 22 may apply upsampling using linear interpolation and transposed convolution to increase the temporal resolution of weekly features. On the other hand, when integrating with monthly features, the scale matching unit 22 may apply downsampling using average pooling and convolution to aggregate information and reduce the temporal resolution of weekly features.

[0045] For example, on a monthly temporal scale, the temporal resolution of data is low, so it is necessary to increase the temporal resolution (upsampling) when integrating with a weekly or daily temporal scale. For example, the scale matching unit 22 can convert monthly features to a finer temporal scale by applying linear interpolation, spline interpolation, and a transposed convolution layer. Further, the scale matching unit 22 can also perform more accurate interpolation in the time direction by using a recurrent neural network (RNN) or an autoregressive model, for example.

[0046] (Learning Unit 23) The learning unit 23 applies residual connections to integrate a plurality of features extracted at different temporal scales. The learning unit 23 can use, but is not limited to, residual connections based on a short-term scale, residual connections based on a long-term scale, residual connections that treat all scales equally, and residual connections introducing learnable weights. Hereinafter, an example of the learning unit 23 having a typical residual connection configuration will be described.

[0047] For example, the learning unit 23 may perform residual connection based on a short-term (daily) temporal scale. The learning unit 23 can further emphasize abrupt short-term changes in time-series data by integrating the output of the intermediate processing layer 231 and the output of the feature extraction unit 21 using the adder 232 based on daily short-term features. Specifically, the learning unit 23 may integrate daily features by auxiliary adding weekly and monthly features thereto. For example, the learning unit 23 may use daily features as main components, temporally scale the weekly and monthly features, and then add the scaled features. Accordingly, the learning unit 23 can output the prediction result 35 that appropriately reflects the influence of a long-term trend while retaining short-term fluctuations of the time-series data 31.

[0048] In addition, the learning unit 23 may perform residual connection based on a long-term (weekly or monthly) temporal scale. The learning unit 23 can emphasize a more overall trend of time-series data by integrating the output of the intermediate processing layer 231 and the output of the feature extraction unit 21 using the adder 232 based on long-term features (weekly or monthly). For example, the adder 232 can strongly reflect long-term fluctuation patterns in the prediction result 35 by using weekly or monthly features as main components and adding daily features as correction information. Through such processing, the learning unit 23 can suppress the influence of short-term noise and achieve stable prediction.

[0049] Furthermore, the learning unit 23 may perform residual connections to treat each time scale equally. For example, the learning unit 23 may add or average the daily, weekly, and monthly features to treat all time scales equally, and then integrate the output of the intermediate processing layer 231 and the output of the feature extraction unit 21 in the adder 232. Specifically, the learning unit 23 may integrate the features of each time scale by adding them with the same weight or by taking a simple average, and output the prediction result 35. This allows the learning unit 23 to obtain a prediction result 35 that makes balanced use of the overall information without being biased towards a particular time scale.

[0050] Furthermore, the learning unit 23 may perform residual connections so that the weights between time scales themselves can be learned. For example, the learning unit 23 may assign learnable weights for each time scale, allowing the model to dynamically learn the importance between time scales, and the output of the intermediate processing layer 231 and the output of the feature extraction unit 21 may be integrated in the adder 232. For example, the adder 232 of the learning unit 23 may multiply the features of each time scale by a learnable scalar coefficient and then add them, allowing the model to automatically adjust the optimal relationship between time scales. With such a configuration, the short-term and long-term effects can be flexibly changed according to the time series data 31, enabling more adaptive predictions based on the properties of the time series data 31.

[0051] (Example of Operation) Figure 7 is a flowchart showing an example of the operation of the information processing device 10. Figure 7 shows an example of a process for predicting future temporal fluctuations of time series data from time series data represented by multiple temporal scales. The operation of the information processing device 10 described with reference to Figure 7 may correspond to one of the information processing methods. The operation of each step in Figure 7 may be executed based on control by the control unit 11 of the information processing device 10.

[0052] In step S1, the control unit 11 acquires time series data (first time series data and second time series data) represented on multiple time scales. For example, the control unit 11 may acquire daily data 321, weekly data 322, and monthly data 323. The control unit 11 may acquire daily data 321, weekly data 322, and monthly data 323 by dividing the time series data 31, or it may acquire daily data 321, weekly data 322, and monthly data 323 that have been previously divided from the same time series data 31.

[0053] In step S2, the control unit 11 extracts features from each time series data acquired in step S1. That is, the control unit 11 extracts features from the first time series data to obtain the first features, and extracts features from the second time series data to obtain the second features. Specifically, the control unit 11 may input each time series data to the feature extraction unit 32 to obtain features from each time series data. As mentioned above, the feature extraction unit 21 may be a pre-trained model such as a neural network.

[0054] In step S3, the control unit 11 matches the features of each time-series data acquired in step S2. For example, the control unit 11 may convert the first feature to match the temporal scale of the second feature, or convert the second feature to match the temporal scale of the first feature.

[0055] In step S4, the control unit 11 predicts the future temporal variation of the time series data based on the characteristics of each time series data after the temporal scale has been aligned in step S3 and the characteristics of each time series data before the temporal scale has been aligned. For example, the control unit 11 may input the first and second characteristics after the temporal scale has been aligned and the first and second characteristics before the temporal scale has been aligned into a trained model to obtain a prediction result of the future temporal variation of the time series data. Specifically, the control unit 11 may input the first and second characteristics after the temporal scale has been aligned and the first and second characteristics before the temporal scale has been aligned into a pre-trained learning unit 23 to obtain a prediction result.

[0056] In step S5, the control unit 11 outputs the prediction result obtained in step S4. For example, the control unit 11 may store the prediction result in the storage unit 12 or display it on the display of the output unit 15. After completing the processing in step S5, the control unit 11 terminates the processing of the flowchart.

[0057] As described above, the information processing device 10 acquires a first feature that shows the characteristics of the time series data at a first temporal scale, and a second feature that shows the characteristics at a second temporal scale, which represent values ​​corresponding to the passage of time. The information processing device 10 aligns the temporal scales of the first and second features. The information processing device 10 inputs the first and second features after the temporal scale alignment, and the first and second features before the temporal scale alignment, into a trained model to predict future temporal fluctuations of the time series data.

[0058] In this way, the information processing device 10 predicts future temporal fluctuations of time series data using multiple features before and after the temporal scales are aligned. Therefore, according to the information processing device 10, even if information loss occurs when aligning temporal scales between features with different temporal scales, it is possible to make highly accurate long-term predictions that accurately capture both short-term fluctuations and long-term trends.

[0059] The information processing device 10 may acquire a first time series data, which is time series data represented on a first time scale, and a second time series data, which is time series data represented on a second time scale. The information processing device 10 may extract features from the first time series data to acquire a first feature. The information processing device 10 may extract features from the second time series data to acquire a second feature. In this way, by acquiring a first feature from the first time series data and a second feature from the second time series data, a prediction result 35 can be obtained based on features corresponding to each time scale using a common feature extraction unit 21.

[0060] The information processing device 10 may also include a feature extraction unit 21, a scale matching unit 22, an intermediate processing layer 231, and a post-processing layer 233. The feature extraction unit 21 may be a feature extraction unit that inputs first time series data and outputs a first feature, and inputs second time series data and outputs a second feature. The scale matching unit 22 may input the first and second features output from the feature extraction unit 21 and output the first and second features after their temporal scales have been matched. The intermediate processing layer 231 may input the first and second features output from the scale matching unit 22 after their temporal scales have been matched and output the first and second features after intermediate processing has been performed. The post-processing layer 233 may input the first and second features after intermediate processing output from the intermediate processing layer 231 and the first and second features output from the feature extraction unit 21 and output a prediction result showing the future temporal fluctuations of the time series data. Here, the trained model for predicting future temporal variations in time series data may be the models of the intermediate processing layer 231 and the post-processing layer 233 that have been pre-trained using training data.

[0061] With this configuration, the post-processing layer 233 receives not only the output from the intermediate processing layer 231, but also the output from the feature extraction unit 21, skipping the scale matching unit 22 and the intermediate processing layer 231. Therefore, the information processing device 10 can perform long-term predictions with high accuracy that reflect both short-term fluctuations and long-term trends.

[0062] In this embodiment, the information processing device 10 has been described primarily as an entity that inputs time-series data of multiple time scales into a feature extraction unit 21 having a common configuration to extract features of a given time scale. However, the device is not limited to this configuration. For example, the information processing device 10 may input the same time-series data 31 into multiple configurations that output features corresponding to predetermined time scales, thereby acquiring features of multiple time scales.

[0063] As described above, the information processing device 10 predicts time series data using a model that employs multiscale residual connections. Specifically, the information processing device 10 size-matches multiple features with different temporal scales for each temporal scale, and then uses residual connections to predict future time series data using a pre-trained model that has been pre-learned to capture short-term fluctuations and long-term trends. Therefore, the information processing device 10 can prevent information loss between temporal scales when integrating features from different temporal scales (daily, weekly, monthly, etc.), enabling long-term predictions with high accuracy.

[0064] This disclosure is not limited to the embodiments described above. For example, multiple blocks shown in the block diagram may be combined, or a single block may be divided. Multiple steps shown in the flowchart may be performed in parallel or in a different order, depending on the processing capacity of the device performing each step, or as necessary, instead of being performed in chronological order as described. Other modifications are possible without departing from the spirit of this disclosure.

[0065] The following additional information is disclosed regarding the embodiments described above.

[0066] [Note 1] An information processing device comprising a control unit that acquires a first feature representing the characteristics of time series data on a first temporal scale and a second feature representing the characteristics on a second temporal scale, aligns the temporal scales of the first and second features, and inputs the first and second features after the temporal scale alignment and the first and second features before the temporal scale alignment into a trained model to predict future temporal fluctuations of the time series data.

[0067] [Addendum 2] The information processing apparatus according to Addendum 1, wherein the control unit acquires first time series data which is time series data represented on a first time scale and second time series data which is time series data represented on a second time scale, extracts the features of the first time series data from the first time series data to acquire the first features, and extracts the features of the second time series data from the second time series data to acquire the second features.

[0068] [Note 3] The information processing apparatus according to Note 2, comprising: a feature extraction unit that inputs the first time series data and outputs the first feature, and inputs the second time series data and outputs the second feature; a scale matching unit that inputs the first and second features output from the feature extraction unit and outputs the first and second features after the temporal scale has been matched; an intermediate processing layer that inputs the first and second features output from the scale matching unit and outputs the first and second features after intermediate processing has been performed; and a post-processing layer that inputs the first and second features output from the intermediate processing layer and the first and second features output from the feature extraction unit and outputs a prediction result showing the future temporal variation of the time series data, wherein the trained model is a model of the intermediate processing layer and the post-processing layer that has been trained in advance using training data.

[0069] [Appendix 4] An information processing method using an information processing device, comprising: acquiring a first feature that shows the characteristics of time series data on a first temporal scale and a second feature that shows the characteristics on a second temporal scale; aligning the temporal scales of the first feature and the second feature; and inputting the first feature and the second feature after the temporal scales have been aligned and the first feature and the second feature before the temporal scales have been aligned into a trained model to predict future temporal fluctuations of the time series data.

[0070] 10: Information processing unit 11: Control unit 12: Memory unit 13: Communication unit 14: Input unit 15: Output unit 21: Feature extraction unit 22: Scale matching unit 23: Learning unit 231: Intermediate processing layer 232: Adder 233: Post-processing layer 31: Time series data 321: Daily data 322: Weekly data 323: Monthly data 35: Prediction results

Claims

1. An information processing device comprising a control unit that acquires a first feature representing the characteristics of a first temporal scale and a second feature representing the characteristics of a second temporal scale from time series data showing values ​​according to the passage of time; aligns the temporal scales of the first and second features; inputs the first and second features after the temporal scale alignment and the first and second features before the temporal scale alignment into a trained model to predict future temporal fluctuations of the time series data.

2. The information processing apparatus according to claim 1, wherein the control unit acquires first time series data which is time series data represented on a first time scale and second time series data which is time series data represented on a second time scale, extracts features of the first time series data from the first time series data to acquire the first features, and extracts features of the second time series data from the second time series data to acquire the second features.

3. An information processing apparatus according to claim 2, comprising: a feature extraction unit that inputs the first time series data and outputs the first feature, and inputs the second time series data and outputs the second feature; a scale matching unit that inputs the first and second features output from the feature extraction unit and outputs the first and second features after the temporal scale has been matched; an intermediate processing layer that inputs the first and second features output from the scale matching unit and outputs the first and second features after intermediate processing has been performed; and a post-processing layer that inputs the first and second features output from the intermediate processing layer and the first and second features output from the feature extraction unit and outputs a prediction result showing the future temporal variation of the time series data, wherein the trained model is a model of the intermediate processing layer and the post-processing layer that has been trained in advance using training data.

4. An information processing method using an information processing device, comprising: acquiring a first feature indicating characteristics on a first temporal scale and a second feature indicating characteristics on a second temporal scale of time series data showing values ​​corresponding to the passage of time; aligning the temporal scales of the first feature and the second feature; and inputting the first feature and the second feature after the temporal scales have been aligned and the first feature and the second feature before the temporal scales have been aligned into a trained model to predict future temporal fluctuations of the time series data.