Program, method, information processing apparatus, and system

By combining first and second time-series data with similar progression in both forward and backward directions, the program extends the learning data time range, enhancing the quality of time-series prediction models.

JP2025078994AActive Publication Date: 2025-05-21AI INSIDE INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023191370
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-09
Publication Date
2025-05-21
Estimated Expiration
2043-11-09

AI Technical Summary

Technical Problem

The period range of learning data cannot be extended, limiting the effectiveness of time-series prediction models.

Method used

A program executed by a computer that acquires first time-series data, identifies second data with similar time-series progression from multiple data sources, combines this data with the first data in both forward and backward time directions to create combined data, and trains a learning model using this combined data.

Benefits of technology

This approach extends the time range of learning data, enabling the creation of higher-quality time-series prediction models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025078994000001_ABST
    Figure 2025078994000001_ABST
Patent Text Reader

Abstract

To extend the time range of training data.SOLUTION: A program to be executed by a computer including a processor and a memory causes the processor to execute: a first data acquisition step of acquiring first data that is time-series data; a second data acquisition step of acquiring second data from a plurality of time-series data, the second data having a time-series trend similar to that of the first data, based on the first data acquired in the first data acquisition step; a concatenation step of generating concatenated data by concatenating the second data to at least one of preceding and succeeding positions of the first data in the time direction; and a training step of training a learning model based on the concatenated data generated in the concatenation step.SELECTED DRAWING: Figure 11
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a program, a method, an information processing apparatus, and a system.

Background Art

[0002] A technique for learning a time-series prediction model based on time-series data is known. Patent Document 1 discloses that it is possible to reduce the man-hours required for selecting data to be used in a transition destination model from among a plurality of transition source data, and to appropriately determine whether or not the transition source model can be transitioned. Patent Document 2 discloses that an optimal learning model for each worker is generated according to the tendency of the collected data for each worker.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0004] There is a problem that the period range of learning data cannot be extended. Therefore, the present disclosure has been made to solve the above problems, and an object thereof is to provide a technique for extending the period range of learning data.

Means for Solving the Problems

[0005] A program to be executed by a computer having a processor and a memory unit, the program executing the following steps: a first data acquisition step in which the processor acquires first data, which is time series data; a second data acquisition step in which, based on the first data acquired in the first data acquisition step, second data, the time series progression of which is similar to that of the first data, is acquired from multiple time series data; a combination step in which the second data is combined with at least one of the first data in the forward and backward time direction to create combined data; and a learning step in which a learning model is trained based on the combined data created in the combination step. Effect of the Invention

[0006] According to the present disclosure, the time range of the learning data can be extended. [Brief description of the drawings]

[0007] [Figure 1] FIG. 2 is a block diagram showing the functional configuration of the system 1. [Diagram 2] FIG. 2 is a block diagram showing the functional configuration of the server 10. [Diagram 3] 2 is a block diagram showing the functional configuration of a user terminal 20. FIG. [Figure 4] FIG. 13 is a diagram showing the data structure of a user table 1012. [Diagram 5] FIG. 13 is a diagram showing the data structure of a main table 1013. [Figure 6] 13 is a diagram showing the data structure of an auxiliary table 1014. FIG. [Figure 7] FIG. 13 is a diagram showing the data structure of a candidate table 1015. [Figure 8] FIG. 13 is a diagram showing the data structure of a model table 1021. [Figure 9] 13 is a flowchart showing the operation of a data extension process. [Figure 10] 13 is a flowchart showing an operation of an initial parameter setting process. [Figure 11] FIG. 1 is a first conceptual diagram illustrating the concept of data extension processing. [Figure 12] FIG. 11 is a second conceptual diagram illustrating the concept of the data extension process. [Figure 13] FIG. 2 is a block diagram showing the basic hardware configuration of a computer 90. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0008] Hereinafter, an embodiment of the present disclosure will be described with reference to the drawings. In all the drawings explaining the embodiment, the same reference numerals are given to common components, and repeated explanations are omitted. Note that the following embodiment does not unduly limit the contents of the present disclosure described in the claims. In addition, not all of the components shown in the embodiment are essential components of the present disclosure. In addition, each figure is a schematic diagram and is not necessarily illustrated strictly.

[0009] <System 1 Configuration> The system 1 in the present disclosure is an information processing system capable of providing an information processing service for learning a time series prediction model based on time series data. The system 1 includes an information processing device, a server 10, and a user terminal 20, which are connected via a network N. FIG. 1 is a block diagram showing the functional configuration of the system 1. FIG. 2 is a block diagram showing the functional configuration of the server 10. As shown in FIG. FIG. 3 is a block diagram showing the functional configuration of the user terminal 20. As shown in FIG.

[0010] Each information processing device is configured by a computer equipped with a calculation device and a storage device. The basic hardware configuration of the computer and the basic functional configuration of the computer realized by the hardware configuration will be described later. For the server 10 and the user terminal 20, descriptions that overlap with the basic hardware configuration and basic functional configuration of the computer described later will be omitted.

[0011] <Server 10 Configuration> The server 10 is an information processing device that provides an information processing service for learning a time series prediction model based on time series data. The server 10 includes a storage unit 101 and a control unit 104 .

[0012] <Configuration of the storage unit 101 of the server 10> The storage unit 101 of the server 10 includes an application program 1011 , a user table 1012 , a main table 1013 , an auxiliary table 1014 , a candidate table 1015 , and a model table 1021 .

[0013] The application program 1011 is a program for causing the control unit 104 of the server 10 to function as each functional unit. Application programs 1011 include applications such as a web browser application.

[0014] User table 1012 is a table that stores and manages information about member users (hereinafter, users) who use the service. When a user registers to use the service, the user's information is stored in a new record in user table 1012. This allows the user to use the service according to the present disclosure. The user table 1012 is a table having columns of user ID and user name, with the user ID as the primary key. FIG. 4 is a diagram showing the data structure of the user table 1012. As shown in FIG.

[0015] The user ID is an item for storing user identification information for identifying a user. The user identification information is an item for which a unique value is set for each user. The user name is an item for storing the name of the user. The user name may be set to any character string such as a nickname instead of a name.

[0016] The main table 1013 is a table for storing and managing information relating to the main data (main data information). The main table 1013 is a table having a main data ID as a primary key, and columns of a main data ID, a user ID, main data, and attribute data. FIG. 5 is a diagram showing the data structure of the main table 1013.

[0017] The main data ID is an item for storing main data identification information for identifying the main data, and a unique value is stored for each main data. The user ID is an item for storing user identification information for identifying a user. Primary data is an item that stores time series data used to train a learning model (time series prediction model). Time series data represents data points ordered over a series of time points or time intervals. It consists of data collected at specific time intervals (e.g. daily, weekly, monthly, etc.). Time series data is characterized by the fact that each data point is consecutive in time, and the order of the data forms the meaning of the data. For example, the primary data may include stock price time series data indicating the closing prices of a stock over a particular period of time, the data being collected on a daily, weekly, monthly, etc. basis. The primary data also includes time series data on the price movements of automobiles and other goods, including data tracking changes in the average selling price of particular models or brands of automobiles over time. In addition, the time series data in the present disclosure does not necessarily have to be time series data of one series (index), but includes a data set including multiple series (indexes). For example, time series data of stock trading is a data set including multiple series (indexes) such as time series data of stock prices and stock price volume (amount of stocks bought and sold). Such a data set is also included in the time series data in the present disclosure. Attribute data is an item that stores metadata about the contents of the main data. Attribute data is information that is used to understand, interpret, and use the main data. Specifically, attribute data includes the following information: Type of data: Basic information about what the time series data represents, for example temperature, humidity, stock prices, sales, etc. · Data origin: information indicating which organization or institution provides the data or by what means or method the data was collected. Collection Frequency: Information indicating whether data is collected daily, hourly, or another frequency. Geographic information: if the data relates to a specific place or region, information about that place or region. Units: The units in which the data values ​​are expressed. For example, Celsius or Fahrenheit for temperature, or currency units for stock prices.

[0018] The auxiliary table 1014 is a table for storing and managing information relating to auxiliary data (auxiliary data information). The auxiliary table 1014 is a table having the auxiliary data ID as a primary key, and columns of the auxiliary data ID, user ID, auxiliary data, and attribute data. FIG. 6 is a diagram showing the data structure of the auxiliary table 1014. As shown in FIG.

[0019] The auxiliary data ID is an item for storing auxiliary data identification information for identifying auxiliary data, and a unique value is stored for each auxiliary data. The user ID is an item for storing user identification information for identifying a user. The auxiliary data is an item that stores time-series data used to supplement or extend the main data. Note that the time-series data is similar to the main data item of the main table 1013. Auxiliary data includes any data used to improve the quality of a machine learning model, improve analysis accuracy, or extend the coverage of the main data when training a machine learning model using the main data. Ancillary data includes data obtained from a data source different from the primary data. Ancillary data includes information on government statistics, such as demographics, economic indicators, and health information, collected and published by government agencies and government-related organizations. The auxiliary data may be obtained via an API (Application Programming Interface) provided by an external information service, a platform service, etc., or may be automatically collected from any web page using a technique such as scraping. The attribute data is an item that stores metadata related to the contents of the auxiliary data. The explanation of the metadata is the same as that of the attribute data item of the main table 1013, so the explanation will be omitted.

[0020] The candidate table 1015 is a table for storing and managing information related to candidates (candidate information). The candidate table 1015 is a table having a candidate data ID as a primary key, and columns of candidate data ID, auxiliary data ID, candidate data, and extraction conditions. FIG. 7 is a diagram showing the data structure of the candidate table 1015. As shown in FIG.

[0021] The candidate data ID is an item for storing candidate data identification information for identifying candidate data. A unique value is stored for each candidate data as the candidate data ID. The auxiliary data ID is an item for storing auxiliary data identification information for identifying auxiliary data. The candidate data is a part of the auxiliary data, and is an item that stores time-series data used to complement or extend the main data. Note that the time-series data is similar to the main data item of the main table 1013. The extraction condition is an item for storing the extraction condition when extracting candidate data from the auxiliary data identified by the auxiliary data identification information. For example, the extraction condition stores information about the start position (start line) and end position (end line) when extracting candidate data from the auxiliary data.

[0022] The model table 1021 is a table for storing and managing information about a learning model (learning model information). The model table 1021 is a table having a model ID as a primary key, and columns of a model ID, a learning model, an initial parameter, a post-learning parameter, a main data ID, a candidate data ID, and an extension condition. FIG. 8 is a diagram showing the data structure of the model table 1021. As shown in FIG.

[0023] The model ID is an item that stores model identification information for identifying a learning model. A learning model is an item that stores a learning model related to a time series prediction model. The time series prediction model includes any time series deep learning model such as a recurrent neural network (RNN), a long short-term memory (LSTM), or a gated recurrent unit (GRU). A time series forecasting model is an inference model that takes time series data as input data and outputs (infers) future indices. For example, a time series forecasting model is an inference model that takes past sales data, weather data, and other data as input data and outputs (infers) future sales figures and temperature indices. The input data may include information regarding seasonality and specific events (such as sales). The output data may include information regarding the confidence levels and ranges of the distribution. The learning process of the time series prediction model will be described later. A time series forecasting model is, for example, a type of machine learning, artificial intelligence, or deep learning model. The time series prediction model does not need to be a single learning model, but may be realized by switching between multiple independent learning models for each product category or regional information. As an example of a time series prediction model, a deep learning model using a deep neural network in deep learning will be described. The time series prediction model does not necessarily have to be a deep learning model, and may be any machine learning or artificial intelligence model. Future market trends and temperature trends are estimated by applying a time series prediction model to product sales history and temperature fluctuation information as input data. In other words, users of the service according to the present disclosure can estimate future sales and temperatures without actually visiting a brick-and-mortar store or looking up weather information. The initial parameters are items that store parameters at the start of learning when learning a learning model. For example, in deep learning models, the quality and speed of learning generally depend heavily on the selection of initial parameters. If inappropriate initial parameters are selected, learning may become slow or converge to a local optimum. The initial parameters may be stored in association with indices indicating the quality of the learning model (various errors, accuracy, matching rate, etc.) or indices indicating the quality of the learning process (convergence speed, learning curve, etc.).The initial parameters may be stored in association with information such as indices indicating the superiority or inferiority of the initial parameters (priority). The trained parameters are the parameters that have been optimized through the training process. The initial parameters are adjusted to minimize the loss function. The main data ID is an item for storing main data identification information of the main data used as learning data when learning a learning model. The candidate data ID is an item for storing candidate data identification information of the candidate data used as learning data when learning the learning model. In the present disclosure, the main data is subjected to data augmentation by the candidate data and used for learning the learning model. The extension condition is an item for storing an extension condition when the main data is extended by the candidate data. Specifically, the candidate data is combined with the main data based on the information stored in the extension condition.

[0024] <Configuration of the control unit 104 of the server 10> The control unit 104 of the server 10 includes a user registration control unit 1041 and a learning unit 1042. The control unit 104 executes an application program 1011 stored in the storage unit 101, thereby realizing each functional unit.

[0025] The user registration control unit 1041 performs processing to store, in the user table 1012, information on users who wish to use the service according to the present disclosure. The information stored in the user table 1012 is generated when a user opens a web page operated by a service provider from any information processing terminal, enters information into a specific input form, and transmits the information to the server 10. The user registration control unit 1041 stores the received information in a new record in the user table 1012, completing the user registration. This allows the user stored in the user table 1012 to use the service. Before the user registration control unit 1041 registers user information in the user table 1012, the service provider may carry out a predetermined examination to restrict whether or not the user is permitted to use the service. The user ID may be any character string or number that can identify the user, any character string or number desired by the user, or an arbitrary character string or number may be automatically set by the user registration control unit 1041.

[0026] The learning unit 1042 executes a learning process. In the present disclosure, the learning unit 1042 can execute a learning step of learning a learning model based on the combined data created in the combining step.

[0027] <Configuration of User Terminal 20> The user terminal 20 is an information processing device operated by a user who uses a service. The user terminal 20 may be, for example, a mobile terminal such as a smartphone or a tablet, a stationary PC (Personal Computer), or a laptop PC. It may also be a wearable terminal such as an HMD (Head Mount Display) or a wristwatch-type terminal. The user terminal 20 includes a storage unit 201 , a control unit 204 , an input device 206 , and an output device 208 .

[0028] <Configuration of the storage unit 201 of the user terminal 20> The storage unit 201 of the user terminal 20 includes a user ID 2011 and an application program 2012 .

[0029] The user ID 2011 is an account ID of the user. The user transmits the user ID 2011 from the user terminal 20 to the server 10. The server 10 identifies the user based on the user ID 2011 and provides the user with the service according to the present disclosure. The user ID 2011 includes information such as a session ID temporarily assigned by the server 10 when identifying the user who is using the user terminal 20.

[0030] The application program 2012 may be stored in advance in the storage unit 201, or may be configured to be downloaded from a web server operated by a service provider via a communication IF. The application programs 2012 include applications such as a web browser application. The application program 2012 includes an interpreted programming language such as JavaScript™ that runs on a web browser application stored on the user terminal 20 .

[0031] <Configuration of the control unit 204 of the user terminal 20> The control unit 204 of the user terminal 20 includes an input control unit 2041 and an output control unit 2042. The control unit 204 executes an application program 2012 stored in the storage unit 201, thereby realizing each functional unit.

[0032] <Configuration of the input device 206 of the user terminal 20> The input device 206 of the user terminal 20 includes a camera 2061 , a microphone 2062 , a position information sensor 2063 , a motion sensor 2064 , and a touch device 2065 .

[0033] <Configuration of the output device 208 of the user terminal 20> The output device 208 of the user terminal 20 includes a display 2081 and a speaker 2082 .

[0034] <System 1 Operation> Each process of the system 1 will be described below. FIG. 9 is a flowchart showing the operation of the data extension process. FIG. 10 is a flowchart showing the operation of the initial parameter setting process. FIG. 11 is a first conceptual diagram illustrating the concept of the data extension process. FIG. 12 is a second conceptual diagram illustrating the concept of the data extension process.

[0035] <Data expansion processing> The data extension process is a process for extending the main data with auxiliary data.

[0036] <Outline of data augmentation processing> The data expansion process is a series of processes that accepts the selection of main data to be expanded, selects auxiliary data to be used for expanding the main data, extracts one or more candidate data from the auxiliary data, selects a specific candidate data from the one or more candidate data, and creates combined data by combining the main data and the candidate data. In the present disclosure, the first data, the second data, and the third data are time-series data having one or more series.

[0037] <Details of data augmentation process> The data extension process will be described in detail below.

[0038] In step S101, the control unit 104 of the server 10 executes a first data acquisition step of acquiring first data, which is time-series data. Specifically, the user operates the input device 206 of the user terminal 20 to execute a browser application or the like, and opens the data extension page D1 by inputting the URL of a web page (data extension page) for executing the data extension process. The control unit 204 of the user terminal 20 transmits a request including a user ID 2011 for opening the data extension page to the server 10.

[0039] When the server 10 receives the request, it generates a data expansion page and transmits it to the user terminal 20. The control unit 204 of the user terminal 20 displays the data expansion page on the display 2081 of the user terminal 20 for presentation. The control unit 104 of the server 10 transmits one or more pieces of main data information stored in the main table 1013 to the user terminal 20, and the control unit 204 of the user terminal 20 may list the one or more pieces of main data information on the data expansion page in a selectable manner based on the received one or more pieces of main data information. Similarly, the control unit 104 of the server 10 may transmit one or more pieces of auxiliary data information stored in the auxiliary table 1014 to the user terminal 20, and the control unit 204 of the user terminal 20 may list one or more pieces of auxiliary data on the data expansion page in a selectable manner based on the received one or more pieces of auxiliary data.

[0040] The user selects the main data listed on the data expansion page by operating the input device 206 of the user terminal 20. The control unit 204 of the user terminal 20 transmits the main data ID of the selected main data to the server 10. The control unit 104 of the server 10 receives the main data ID. The control unit 104 of the server 10 searches the item of the main data ID in the main table 1013 based on the main data ID, and acquires and accepts the main data (first data). Note that the control unit 104 of the server 10 may acquire and accept a plurality of main data (first data).

[0041] In step S102, the control unit 104 of the server 10 executes a third data acquisition step to acquire third data, which is time series data having a third period range longer than the first period range of the first data acquired in the first data acquisition step. Specifically, the user selects auxiliary data listed on the data expansion page by operating the input device 206 of the user terminal 20. The control unit 204 of the user terminal 20 transmits the auxiliary data ID of the selected auxiliary data to the server 10. The control unit 104 of the server 10 receives the auxiliary data ID. The control unit 104 of the server 10 searches the auxiliary data ID item of the auxiliary table 1014 based on the auxiliary data ID, and obtains and accepts the auxiliary data (third data).

[0042] The control unit 104 of the server 10 may acquire the auxiliary data without accepting a selection operation from the user. For example, the control unit 104 of the server 10 may search, acquire, and accept third data suitable for extending the first data from the auxiliary table 1014 based on information such as the first data acquired in step S101 and metadata (attribute data) stored in association with the first data. The control unit 104 of the server 10 may be configured to acquire and accept all or a part of any auxiliary data stored in the auxiliary table 1014. The control unit 104 of the server 10 may acquire and accept a plurality of auxiliary data (third data). The control unit 104 of the server 10 may identify the third data by using a machine learning model, a deep learning model, or any other artificial intelligence model that uses information such as the first data and metadata (attribute data) stored in association with the first data as input data and outputs information for identifying the third data (third data ID).

[0043] In step S103, the control unit 104 of the server 10 executes a candidate extraction step of extracting a plurality of candidate data, which are a portion of the third data acquired in the third data acquisition step and are a plurality of time series data included in the third period range. Specifically, the control unit 104 of the server 10 extracts, from the acquired auxiliary data, one or more candidate data included in a part of the time range of the auxiliary data. Specifically, the control unit 104 of the server 10 may extract as candidate data a part of the time range of the acquired auxiliary data, or may exclude as candidate data a part of the time range of the auxiliary data. The control unit 104 of the server 10 may extract some of the multiple series contained in the auxiliary data as candidate data, or may exclude some of the series and extract them as candidate data.

[0044] In step S103, the candidate extracting step executes a step of extracting a plurality of candidate data having substantially the same period range as the first period range. Specifically, it is preferable that the control unit 104 of the server 10 extracts one or more candidate data having approximately the same period range as the period range (first period range) of the main data selected in step S101 from the acquired auxiliary data. Even if the period range of the main data and the candidate data is substantially the same, the number of data in the time direction may differ between the main data and the candidate data (auxiliary data). In other words, the density of the number of data in the time direction (the number of data per unit time) between the main data and the candidate data (auxiliary data) may differ. In this case, the control unit 104 of the server 10 applies an arbitrary complementation process to the main data and the candidate data (auxiliary data) to make the number of data between the main data and the candidate data (auxiliary data) uniform. It is preferable to apply the complementation process to the candidate data (auxiliary data) rather than to the main data. Imputation processing is a method for filling in missing values ​​or insufficient data between multiple consecutive time series data, and various methods are known. Any imputation processing can be applied, such as forward imputation, backward imputation, linear imputation, mean imputation, median imputation, nearest neighbor imputation, imputation using a machine learning model, deep learning model, etc.

[0045] In step S103, the candidate extraction step includes a step of extracting first candidate data within a predetermined period range shorter than the third period range, and a step of extracting multiple candidate data included in the predetermined period range from the first candidate data by sequentially shifting the third data forward or backward by a predetermined period in the time direction. Specifically, the control unit 104 of the server 10 extracts a period range from the oldest position (start position) in the time direction to the first period range in the time direction from the auxiliary data, and extracts it as the first candidate data. The control unit 104 of the server 10 extracts a period range from a position (second position) shifted by a predetermined period from the start position in the time direction to the first period range in the time direction from the auxiliary data, and extracts it as the second candidate data. The control unit 104 of the server 10 extracts a period range from a position (third position) shifted by a predetermined period from the second position in the time direction to the first period range in the time direction from the auxiliary data, and extracts it as the third candidate data. The control unit 104 of the server 10 extracts a plurality of candidate data by sequentially shifting the extraction start position of the period range for each predetermined period in this way. The period can be any period such as one day, one week, one month, or one year. It should be noted that the candidate data does not have to be extracted from the oldest position in the time direction of the auxiliary data, but may be extracted from the newest position or any position included in the third period range.

[0046] In step S103, the candidate extraction step includes a step of excluding data that is earlier or later in the time direction than the first period range from the third data acquired in the third data acquisition step, and a step of extracting multiple candidate data from the excluded third data. Specifically, the control unit 104 of the server 10 extracts one or more pieces of candidate data so as not to be included in a period range that is later in the time direction than the first period range in the time direction than the first period range. Specifically, the control unit 104 of the server 10 may be configured to exclude a period range that is later in the time direction than the first period range in the time direction in the third period range from the auxiliary data, and extract candidate data from the excluded auxiliary data. When constructing a time series prediction model for predicting events later (future) than the first time range based on the first data, it is not preferable to use, as learning data, data from the third data that is later in time (future) than the time range of the first data, in consideration of causal relationships. The same applies to the case where a time series prediction model for inferring events before (in the past) the first period range is constructed based on the first data. In this case, the control unit 104 of the server 10 extracts one or more pieces of candidate data so as not to be included in a period range before the first period range in the time direction.

[0047] In step S103, the control unit 104 of the server 10 executes a cycle input step of accepting input of a predetermined cycle period from the user. Specifically, the cycle period may be input based on a value input by the user in a cycle period input field or the like provided on the data expansion page. In this case, the candidate extracting step executes a step of extracting a plurality of candidate data based on the periodic period for which the input was accepted in the periodic input step.

[0048] In step S103, the candidate extracting step executes a step of extracting a plurality of candidate data based on the periodic period specified based on the first data, without receiving an input of the periodic period from the user. Specifically, the control unit 104 of the server 10 may automatically identify a suitable periodic period for extracting auxiliary data when expanding the first data, based on information such as the first data acquired in step S101 and the metadata (attribute data) stored in association with the first data, without receiving an input from the user. For example, when the first data is data that varies at a predetermined period, a periodic period shorter or longer than the period of the main component may be identified based on the period of the main component of the first data. Alternatively, it may be identified based on a period determined according to the type, content, etc. of the first data (data affected by periodic factors such as demographic trends and seasonal variation factors). The control unit 104 of the server 10 may identify the periodic period by using a machine learning model, a deep learning model, or any other artificial intelligence model that outputs the periodic period, with information such as the first data and the metadata (attribute data) stored in association with the first data as input data.

[0049] FIG. 11 is a first conceptual diagram for explaining the concept of data expansion processing. The data expansion processing of the first data D100 (data for 3.5 years, 42 rows) in the auxiliary data D110 (data for 20 years, 240 rows) is explained. The control unit 104 of the server 10 excludes the data D112 (data for 2 years, 24 rows) behind the first data D10 in the time direction from the auxiliary data D110 and specifies it as the auxiliary data D111 (data for 18 years, 216 rows). The control unit 104 of the server 10 extracts a period range of 3.5 years, which is the first period range, in the time direction from the latest position D13 (start position) in the time direction of the first candidate data, and specifies it as the first candidate data D121 (3.5 years, 42 rows of data). The control unit 104 of the server 10 extracts a period range of 3.5 years, which is the first period range, in the time direction from a position (second position) shifted in the time direction by a periodic period of 12 months from the position D13, and specifies it as the second candidate data D122 (3.5 years, 42 rows of data). Similarly, the control unit 104 of the server 10 extracts a period range of 3.5 years, which is the first period range, in the time direction from a position (third position) shifted in the time direction by a periodic period of 12 months from the second position, and specifies it as the third candidate data D123 (3.5 years, 42 rows of data). In this way, the control unit 104 of the server 10 can extract 14 candidate data from the auxiliary data D111 by sequentially shifting the period by period.

[0050] In step S104, the control unit 104 of the server 10 executes a second data acquisition step of acquiring second data having a time series transition similar to that of the first data from the plurality of time series data based on the first data acquired in the first data acquisition step. The second data acquisition step executes a step of acquiring second data from the plurality of candidate data according to the similarity between the first data and the plurality of candidate data. The second data acquisition step includes a step of calculating at least one of a Manhattan distance, a Euclidean distance, a cosine similarity, and a correlation coefficient between the first data and a plurality of candidate data, and a step of acquiring the second data according to the similarity based on the calculated distance. Specifically, the control unit 104 of the server 10 calculates distance elements such as the difference (xij-yij) or product (xij*yij) between xi and yij, where xi is the value of the i-th data (data index is i) from the start position of the first data, and yij is the value of the i-th data (data index is i) from the start position of the j-th candidate data (note that it is assumed that the candidate data has already been subjected to the interpolation process). The distance between the first data and the j-th candidate data can be calculated by accumulating the distance elements for all i. The distance can be calculated using distances such as Manhattan distance (L1 norm), Euclidean distance (L2 norm), cosine similarity, and correlation coefficient. The control unit 104 of the server 10 identifies and acquires the candidate data with the smallest distance as the second data among the distances calculated for the plurality of candidate data. It is not necessary to identify the candidate data with the shortest distance as the second data, and may select a predetermined candidate data as the second data from among the plurality of candidate data with a similarity calculated based on the distance (determined by the inverse of the distance, etc.) equal to or greater than a predetermined value. The control unit 104 of the server 10 may also select a plurality of candidate data as a plurality of second data.

[0051] In step S104, the second data acquisition step executes a step of acquiring second data from the multiple candidate data according to the similarity between the time series data of each sequence included in the first data and the time series data of each sequence included in the multiple candidate data. Specifically, when the first data and the candidate data are time series data consisting of multiple series, distances such as the Manhattan distance (L1 norm), Euclidean distance (L2 norm), cosine similarity, and correlation coefficient described above are calculated for each series, and the similarity between the first data and the candidate data is calculated by combining these distances calculated for the series. For example, at least one of the average (average, median, etc.) of the distance or similarity calculated for each sequence, the weighted average, the maximum value, and the minimum value may be determined as the similarity between the first data and the candidate data. The control unit 104 of the server 10 identifies and acquires the candidate data with the highest similarity as the second data among the similarities calculated for the multiple candidate data. Note that it is not necessary to identify the candidate data with the highest similarity as the second data, and for example, a predetermined candidate data may be selected as the second data from multiple candidate data with similarities equal to or greater than a predetermined value. Also, the control unit 104 of the server 10 may select multiple candidate data as multiple second data.

[0052] In step S104, the second data obtaining step includes a step of obtaining second data from the plurality of candidate data in accordance with the number of similar series between the first data and the plurality of candidate data. Specifically, it is assumed that the first data and the candidate data are time series data having three series, the first series, the second series, and the third series. In this case, for the multiple candidate data, the distance (similarity) is calculated for each of the first series, the second series, and the third series. For each of the multiple candidate data, the number of most similar series among the series held by the first data and the multiple candidate data is counted. For example, assume that there are five candidate data: first candidate data, second candidate data, third candidate data, fourth candidate data, and fifth candidate data. For the first candidate data, the number of sequences most similar to the first data for each of the first sequence, second sequence, and third sequence is 0. Similarly, when there are 0 second candidate data, 2 third candidate data, 0 fourth candidate data, and 1 fifth candidate data, the third candidate data, which is the most similar candidate data, is identified as the second data and acquired.

[0053] In step S105, the control unit 104 of the server 10 executes a combining step of generating combined data by combining the second data with at least one of the first data and the second data in the time direction. The combining step may include a step of creating combined data by combining one or more pieces of second data forward in a time direction of the first data. Specifically, the control unit 104 of the server 10 combines the candidate data (second data) selected in step S104 with the main data (first data) selected in step S101 before it in the time direction. Specifically, the rearmost data in the time direction of the second data is combined with the frontmost data in the time direction of the first data, which is time-series data. This makes it possible to create time-series data (combined data) in which the second data and the first data are consecutive in that order in the time direction.

[0054] The combining step may include a step of generating combined data by combining one or more pieces of second data backward in a time direction with the first data. Similarly, the control unit 104 of the server 10 may combine the candidate data (second data) selected in step S104 after the main data (first data) selected in step S101 in the time direction. Specifically, the forwardmost data in the time direction of the second data is combined with the rearmost data in the time direction of the first data, which is time-series data. This makes it possible to create time-series data (combined data) in which the first data and the second data are consecutive in that order in the time direction.

[0055] In step S105, the control unit 104 of the server 10 executes an extension period receiving step of receiving, from the user, an extension period for the first data acquired in the first data acquiring step. The combining step includes a step of combining one or more pieces of second data at least either forward or backward in the time direction of the first data until the period range of the combined data reaches the extended period. Specifically, the user may be configured to be able to input a period (extension period) longer than the first period range for which the first data is desired to be extended into an extension period input field displayed on the data extension page. In this case, the control unit 104 of the server 10 combines one or more second data items with the first data item in the forward time direction according to the extension period received from the user. Specifically, a predetermined number of second data items are combined so that the period range of the combined data becomes the extension period. Note that the period range of the combined data can be set to the extension period by combining ((extension period-first period range)÷first period range) pieces of second data items with the first data item in the forward time direction. In addition, if a predetermined number of second data required to make the combined data into an extended period cannot be selected in step S104, the period range of the first data may be extended by repeatedly combining the second data selected in step S104 with the first data.

[0056] Similarly, the control unit 104 of the server 10 may create combined data having an extended period by combining one or more second data following the first data in the time direction, depending on the extended period received from the user.

[0057] When the period range of the combined data created by combining the first data with one or more second data exceeds the extended period, the control unit 104 of the server 10 may perform processing to make the period range of the combined data the extended period by excluding the exceeded period range of the second data.

[0058] FIG. 12 is a second conceptual diagram illustrating the concept of the data extension process. The control unit 104 of the server 10 creates first combined data D221 (10.5 years, 126 lines of data) for an extended period of 10.5 years by combining the first data D211 (3.5 years, 42 lines of data) which is the main data with two second candidate data D212, D213 (3.5 years, 42 lines of data) extracted from the auxiliary data. Note that, although an example in which the second candidate data D212, D213 are combined with the same data has been described as an example, it is not necessary to combine the same candidate data. Next, the control unit 104 of the server 10 executes a data extension process with the first combined data D221 as new main data (first data). Specifically, with the first combined data D221 as new main data, the A-th candidate data D222 (10.5 years, 126 rows of data), ..., and the Z-th candidate data D229 (10.5 years, 126 rows of data) are extracted from the auxiliary data. Note that the A-th candidate data D222, ..., and the Z-th candidate data D229 may also be combined data created by combining other main data and candidate data. In the present disclosure, the control unit 104 of the server 10 combines 19 pieces of the A-th candidate data D222, ..., and the Z-th candidate data D229 with the first combined data D221 to obtain combined data of 210 years and 2520 rows, which is the extended period, to create second combined data (210 years, 2520 rows of data). In this way, the control unit 104 of the server 10 can create combined data from the main data (first data) and auxiliary data by performing a single data extension process. Furthermore, the control unit 104 of the server 10 can handle the created combined data as main data or auxiliary data and create combined data for any period (extension period) by sequentially executing data extension processes. In this way, even if only time-series data for a limited period of time is available for an event at hand, it is possible to create new, high-quality long-term time-series data by using data augmentation processing. For long-term events, a high-quality learning model can be trained based on the long-term learning data.

[0059] <Initial parameter setting process> The initial parameter setting process is a process for setting initial parameters when learning a learning model.

[0060] <Overview of the initial parameter setting process> The initial parameter setting process is a series of processes that accept the selection of a learning model, acquire learning data for training the learning model, search for initial parameters based on the acquired learning data, and set the initial parameters identified by the search as the initial parameters of the learning model.

[0061] <Details of the initial parameter setting process> The initial parameter setting process will be described in detail below. Prior to the initial parameter setting process, the control unit 104 of the server 10 executes a storage step of storing a plurality of time series prediction models, learning data which is time series data used for learning the plurality of time series prediction models, and learning parameters of the plurality of time series prediction models in association with each other. The storing step is a step of storing a plurality of time series prediction models and optimized parameters of the plurality of time series prediction models in association with each other. The storing step is a step of storing a plurality of time series prediction models and initial parameters of the plurality of time series prediction models in association with each other. Specifically, the control unit 104 of the server 10 executes a learning process based on the initial parameters using the main data or the combined data for one or more learning models stored in the model table 1021, and calculates optimized post-learned parameters. The control unit 104 of the server 10 stores the learning model, the initial parameters, and the post-learned parameters in association with the items of the learning model, the initial parameters, and the post-learned parameters in the model table 1021, respectively.

[0062] In step S301, the control unit of the server 10 executes a learning model selection step of accepting the selection of a learning model. The user operates the input device 206 of the user terminal 20 to execute a browser application or the like, and opens the initial parameter setting page D3 by inputting the URL of a web page (initial parameter setting page) for executing the initial parameter setting process. The control unit 204 of the user terminal 20 transmits a request including a user ID 2011 for opening the initial parameter setting page to the server 10.

[0063] When the server 10 receives the request, it generates an initial parameter setting page and transmits it to the user terminal 20. The control unit 204 of the user terminal 20 displays the initial parameter setting page on the display 2081 of the user terminal 20 for presentation. The control unit 104 of the server 10 transmits one or more pieces of learning model information stored in the model table 1021 to the user terminal 20, and the control unit 204 of the user terminal 20 may list one or more pieces of learning model information on the initial parameter setting page in a selectable manner based on the received one or more pieces of learning model information. The user operates the input device 206 of the user terminal 20 to select a learning model listed on the initial parameter setting page. The control unit 204 of the user terminal 20 transmits the model ID of the selected learning model to the server 10. The control unit 104 of the server 10 receives and accepts the model ID.

[0064] In step S302, the control unit 104 of the server 10 executes a first learning data acquisition step of acquiring first learning data, which is time-series data. Specifically, the control unit 104 of the server 10 searches the model ID item of the model table 021 based on the received model ID, and acquires the main data ID, candidate data ID, and extension condition items. The control unit 104 of the server 10 searches the main data ID item of the main table 1013 based on the acquired main data ID, and acquires the main data. The control unit 104 of the server 10 searches the candidate data ID item of the candidate table 1015 based on the acquired candidate data ID, and acquires the candidate data. In the initial parameter setting process, the first learning data includes any data such as main data, auxiliary data, candidate data, and combined data stored in association with the learning model. For example, the first learning data includes the acquired main data. The first learning data includes combined data obtained by combining the acquired main data and candidate data in step S105 of the data extension process.

[0065] In addition, the control unit 104 of the server 10 may transmit one or more pieces of main data information stored in the main table 1013 to the user terminal 20, and the control unit 204 of the user terminal 20 may list the one or more pieces of main data information on the initial parameter setting page in a selectable manner based on the received one or more pieces of main data information. The user operates the input device 206 of the user terminal 20 to select main data displayed in a list on the initial parameter setting page. The control unit 204 of the user terminal 20 transmits the main data ID of the selected main data to the server 10. The control unit 104 of the server 10 receives the main data ID. The control unit 104 of the server 10 searches the item of the main data ID in the main table 1013 based on the main data ID, and acquires and accepts the main data (first learning data). Note that the control unit 104 of the server 10 may acquire and accept a plurality of main data (first learning data). In this way, the control unit 104 of the server 10 may identify and acquire main data associated with a selected learning model in response to a user's selection of multiple learning models, or may identify and acquire the selected main data in response to a direct selection by the user of multiple main data. Similarly, the control unit 104 of the server 10 may identify and acquire combined data associated with a selected learning model in response to a user's selection of multiple learning models, or may identify and acquire the selected combined data in response to a direct selection by the user of multiple combined data.

[0066] In step S303, the control unit 104 of the server 10 executes a data identifying step of identifying, from the multiple pieces of learning data stored in the storing step, second learning data whose time series progression is similar to that of the first learning data, based on the first learning data acquired in the first learning data acquiring step. The data identification step includes a step of calculating at least one of a Manhattan distance, a Euclidean distance, a cosine similarity, and a correlation coefficient between the first training data and the multiple training data, and a step of identifying the second training data according to the similarity based on the calculated distance. Specifically, the control unit 104 of the server 10 compares the first learning data with at least one of the main data identified based on the main data ID stored in the model table 1021, the candidate data identified based on the candidate data ID, and the combined data obtained by combining the acquired main data and candidate data in step S105 of the data expansion process, and calculates the similarity. When the first learning data is the main data, the candidate data, or the combined data, it is preferable that the comparison target is also the main data, the candidate data, or the combined data. Specifically, the control unit 104 of the server 10 may compare the first learning data with at least one of the main data, the candidate data, and the combined data (hereinafter referred to as target data) having substantially the same period range as the period range of the first learning data to calculate the similarity (the main data is compared with the main data, the candidate data with the candidate data, and the combined data with the combined data). The control unit 104 of the server 10 refers to the model table 1021 to acquire multiple target data stored in association with multiple learning models, initial parameters, and post-learning parameters. Specifically, the control unit 104 of the server 10 calculates distance elements such as the difference (xij-yij) or product (xij*yij) between xi and yij, where xi is the value of the i-th data (data index is i) from the start position of the first learning data, and yij is the value of the i-th data (data index is i) from the start position of the j-th target data (note that it is assumed that the candidate data has already been subjected to the interpolation process). The distance between the first data and the j-th target data can be calculated by accumulating the distance elements for all i. The distance can be calculated using distances such as Manhattan distance (L1 norm), Euclidean distance (L2 norm), cosine similarity, and correlation coefficient. The control unit 104 of the server 10 identifies and acquires the target data with the smallest distance as the second learning data among the distances calculated for the multiple target data. Note that it is not necessary to identify the target data with the shortest distance as the second data, and for example, a predetermined target data may be selected as the second data from multiple target data whose similarity calculated based on the distance (determined by the inverse of the distance, etc.) is equal to or greater than a predetermined value. The control unit 104 of the server 10 may also select multiple target data as multiple second data.

[0067] The data identifying step executes a step of identifying second learning data based on a similarity calculated based on the objective variables of the first learning data acquired in the first learning data acquiring step and the objective variables of the plurality of learning data stored in the storing step. The data identifying step executes a step of not calculating a similarity based on the explanatory variables of the first learning data acquired in the first learning data acquiring step and the explanatory variables of the plurality of learning data stored in the storing step. Specifically, when calculating the distance between the first learning data and one or more target data, the control unit 104 of the server 10 may calculate the distance by considering only the objective variable, without considering the explanatory variables of the first learning data and the target data. In general, explanatory variables of training data have a high-dimensional data structure, whereas objective variables have a one-dimensional or small-dimensional data structure. This allows the second training data to be identified in a shorter processing time and at a lower cost.

[0068] The data identifying step includes a first step of identifying multiple second learning data candidates based on a similarity calculated based on the objective variables of the first learning data acquired in the first learning data acquiring step and the objective variables of the multiple learning data stored in the storing step, and a second step of identifying second learning data based on a similarity calculated based on the explanatory variables of the first learning data acquired in the first learning data acquiring step and the explanatory variables of the multiple second learning data candidates. Specifically, when calculating the distance between the first learning data and one or more target data, the control unit 104 of the server 10 calculates the distance by considering only the objective variable without considering the explanatory variables of the first learning data and the target data, and identifies multiple target data for which the calculated distance is equal to or less than a predetermined value. In other words, the control unit 104 of the server 10 compares the first learning data and the multiple target data with respect to the explanatory variables, and narrows down the multiple target data to relatively similar target data as candidates for the second learning data (second learning data candidates). Next, when calculating the distance between the first learning data and one or more second learning data candidates, the control unit 104 of the server 10 calculates the distance taking into consideration the explanatory variables of the first learning data and the second learning data candidates, and identifies and acquires the second learning data candidate with the smallest distance as the second learning data. Note that the control unit 104 of the server 10 may calculate the distance taking into consideration the objective variables and explanatory variables of the first learning data and the second learning data candidates. In general, explanatory variables of the learning data have a high-dimensional data structure, whereas the objective variable has a one-dimensional or decimal-dimensional data structure. This makes it possible to narrow down the second learning data in a shorter processing time and at a lower cost, while calculating the similarity for a small number of second learning data candidates by taking the explanatory variables into account, thereby making it possible to identify suitable second learning data with high accuracy.

[0069] In step S303, the data identifying step executes a step of identifying second learning data based on the objective variable output by applying the objective variable of the first learning data acquired in the first learning data acquiring step as input data to the second learning model stored in the model storing step. In addition, the control unit 104 of the server 10 can also be configured to identify the second learning data using the objective variables of the first learning data as input data, using a machine learning model, a deep learning model, an artificial intelligence model, or the like that can identify the second learning data using the objective variables as input data. This makes it possible to identify the second training data in a shorter processing time and at a lower cost.

[0070] In the storage step, the control unit 104 of the server 10 executes a parameter acquisition step of acquiring second learning parameters (second optimized parameters, second initial parameters) stored in association with the second learning data identified in the data identification step. Specifically, the control unit 104 of the server 10 acquires at least one of the initial parameters and the learning parameters stored in the model table 1021 in association with the second data (stored in the same record). The control unit 104 of the server 10 stores the acquired second initial parameters and second optimized parameters in the initial parameter and learned parameter fields of the record in the model table 1021 identified based on the model ID selected in step S301.

[0071] In step S304, the control unit 104 of the server 10 executes a learning step in which a learning model is trained based on the first learning data using initial parameters based on the second learning parameters (second optimized parameters or second initial parameters) acquired in the parameter acquisition step. Specifically, the control unit 104 of the server 10 searches the model ID item of the model table 1021 based on the model ID selected in step S301, and obtains the main data, candidate data, extension conditions, initial parameters (second initial parameters), and learning parameters (second optimized parameters). The control unit 104 of the server 10 creates combined data by combining the acquired main data and candidate data according to the process in step S105 of the data extension process.

[0072] The control unit 104 of the server 10 executes a learning process by deep learning for the learning parameters of the deep neural network included in the learning model. The control unit 104 of the server 10 trains the learning model using the combined data as learning data and the second initial parameters or the second optimized parameters as initial parameters. The control unit 104 of the server 10 creates a data set such as training data, test data, and verification data for learning the deep neural network of the learning model based on the combined data. The learning unit 1051 of the server 10 learns the learning parameters of the deep neural network included in the learning model by deep learning based on the created data set.

[0073] In the present disclosure, as an example, in setting initial parameters when training a learning model using combined data, an example has been described in which optimized parameters and initial parameters of a learning model trained using combined data similar to the combined data are used as initial parameters when training the learning model, but the present disclosure is not limited to this. For example, when setting initial parameters when training a learning model using main data, the optimized parameters and initial parameters of a learning model trained using main data similar to the main data may be used as initial parameters when training the learning model.

[0074] In step S304, the control unit 104 of the server 10 executes a correction step of correcting the second learning parameter acquired in the parameter acquisition step based on the similarity between the first learning data acquired in the first learning data acquisition step and the second learning data identified in the data identification step. The modifying step executes a step of modifying the second learning parameter based on a random number within a range determined according to the similarity between the first learning data and the second learning data. Specifically, the control unit 104 of the server 10 may apply processing to the acquired value of the second learning parameter according to the similarity (distance) between the second learning parameter and the first learning data calculated when selecting the candidate data in step S104. Specifically, a random number having a magnitude according to the similarity may be added to or subtracted from the second learning parameter. For example, assuming that the similarity is S, a random number having a range from -S to +S is added to the second learning parameter. When the second learning parameter is a multidimensional quantity, a different random value generated for each dimension may be added to the second learning parameter. As a result, when the first learning data and the second learning data are similar to each other, parameters similar to the second learning parameters are used in the learning process of the first learning data. On the other hand, when the first learning data and the second learning data are not similar to each other, parameters obtained by adding the influence of random numbers to the second learning parameters are used in the learning process of the first learning data. In this way, the second learning parameter is modified according to the degree of similarity between the first learning data and the second learning data, and the contribution ratio of the second learning parameter during the learning process can be controlled according to the degree of similarity between the first learning data and the second learning data.

[0075] The learning step executes a step of learning a learning model using initial parameters based on the second learning parameters corrected in the correction step. The control unit 104 of the server 10 executes a learning process of the learning model using the second learning parameters as initial parameters.

[0076] <Basic computer hardware configuration> 13 is a block diagram showing the basic hardware configuration of a computer 90. The computer 90 includes at least a processor 901, a main storage device 902, an auxiliary storage device 903, and a communication IF 991 (interface). These are electrically connected to each other by a communication bus 921.

[0077] The processor 901 is hardware for executing an instruction set described in a program, and is composed of an arithmetic unit, a register, a peripheral circuit, and the like.

[0078] The main memory device 902 is for temporarily storing programs, data to be processed by the programs, etc. For example, it is a volatile memory such as a DRAM (Dynamic Random Access Memory).

[0079] The auxiliary storage device 903 is a storage device for saving data and programs, such as a flash memory, a hard disk drive (HDD), a magneto-optical disk, a CD-ROM, a DVD-ROM, or a semiconductor memory.

[0080] The communication IF 991 is an interface for inputting and outputting signals for communicating with other computers via a network using a wired or wireless communication standard. The network is composed of the Internet, a LAN, various mobile communication systems constructed by wireless base stations, etc. For example, the network includes 3G, 4G, 5G mobile communication systems, LTE (Long Term Evolution), wireless networks that can connect to the Internet via a specified access point (e.g., Wi-Fi (registered trademark)), etc. In the case of wireless connection, communication protocols include, for example, Z-Wave (registered trademark), ZigBee (registered trademark), Bluetooth (registered trademark), etc. In the case of wired connection, the network also includes a network that is directly connected by a USB (Universal Serial Bus) cable or the like.

[0081] It should be noted that the computer 90 can be virtually realized by distributing all or part of each hardware configuration among multiple computers 90 and connecting them together via a network. In this way, the computer 90 is a concept that includes not only a computer 90 housed in a single housing or case, but also a virtualized computer system.

[0082] <Basic functional configuration of computer 90> A description will now be given of the functional configuration of a computer realized by the basic hardware configuration (FIG. 13) of a computer 90. The computer comprises at least the functional units of a control unit, a storage unit, and a communication unit.

[0083] The functional units of the computer 90 can also be realized by distributing all or part of the functional units among multiple computers 90 connected to each other via a network. The computer 90 is a concept that includes not only a single computer 90 but also a virtualized computer system.

[0084] The control unit is realized by the processor 901 reading out various programs stored in the auxiliary storage device 903, expanding the programs in the main storage device 902, and executing processes according to the programs. The control unit can realize functional units that perform various information processing depending on the type of program. In this way, the computer is realized as an information processing device that performs information processing.

[0085] The storage unit is realized by a main storage device 902 and an auxiliary storage device 903. The storage unit stores data, various programs, and various databases. Furthermore, the processor 901 can secure a storage area corresponding to the storage unit in the main storage device 902 or the auxiliary storage device 903 in accordance with a program. Furthermore, the control unit can cause the processor 901 to execute processes of adding, updating, and deleting data stored in the storage unit in accordance with the various programs.

[0086] The term database refers to a relational database, which is used to manage sets of data called masters and tables in a tabular format structurally defined by rows and columns, by associating them with each other. In a database, a table is called a table or master, a column in a table is called a column, and a row in a table is called a record. In a relational database, relationships between tables and masters can be set and associated. Usually, a column that serves as a primary key for uniquely identifying a record is set in each table and each master, but setting a primary key in a column is not essential. The control unit can cause the processor 901 to add, delete, or update records in a specific table or master stored in the storage unit according to various programs. Furthermore, by storing data, various programs, and various databases in the storage unit, it can be considered that the information processing device and information processing system according to the present disclosure have been manufactured.

[0087] In addition, the database and master in this disclosure may include any data structure (such as a list, a dictionary, an associative array, or an object) in which information is structurally defined. The data structure also includes data that can be considered as a data structure by combining data with a function, class, method, or the like written in any programming language.

[0088] The communication unit is realized by the communication IF 991. The communication unit realizes a function of communicating with other computers 90 via a network. The communication unit can receive information transmitted from other computers 90 and input the information to the control unit. The control unit can cause the processor 901 to execute information processing on the received information in accordance with various programs. In addition, the communication unit can transmit information output from the control unit to other computers 90.

[0089] <Additional Notes> The matters described in the above embodiments will be supplemented below.

[0090] (Appendix 1) A program to be executed by a computer having a processor and a memory unit, the program comprising: a first data acquisition step (S101) in which the processor acquires first data, which is time series data; a second data acquisition step (S104) in which, based on the first data acquired in the first data acquisition step, second data having a time series progression similar to that of the first data from multiple time series data; a combination step (S105) in which the second data is combined with at least one of the first data in the forward and backward time direction to create combined data; and a learning step in which a learning model is trained based on the combined data created in the combination step. This makes it possible to extend the period range of the learning data (combined data) when training a learning model, which is a time series prediction model, from the period range of the first data, thereby creating a learning model with higher quality.

[0091] (Appendix 2) The program according to appended claim 1, wherein the processor executes a third data acquisition step (S102) of acquiring third data, which is time series data having a third period range longer than the first period range of the first data acquired in the first data acquisition step, and a candidate extraction step (S103) of extracting a plurality of candidate data, which is a portion of the third data acquired in the third data acquisition step and is a plurality of time series data included in the third period range, and the second data acquisition step (S104) is a step of acquiring second data from the plurality of candidate data according to the similarity between the first data and the plurality of candidate data. This makes it possible to extend the period range of the learning data (combined data) from the period range of the first data when learning a learning model based on the second data similar to the first data, which is included in a part of a third period that has a longer period range than the first data, thereby creating a learning model of higher quality.

[0092] (Appendix 3) The program according to claim 2, wherein the candidate extracting step (S103) is a step of extracting a plurality of candidate data having substantially the same period range as the first period range. This makes it possible to extend the period range of the learning data (combined data) when learning a learning model based on the second data having substantially the same period range as the first data from the period range of the first data, thereby creating a learning model of higher quality. Note that "substantially the same" refers to the time range of a plurality of candidate data being a time range sufficient to calculate the similarity between the candidate data and the first data, and the time ranges do not necessarily have to be the same.

[0093] (Appendix 4) The program described in Appendix 2, wherein the second data acquisition step (S104) includes a step of calculating at least one of a Manhattan distance, a Euclidean distance, a cosine similarity, and a correlation coefficient between the first data and a plurality of candidate data, and a step of acquiring the second data according to the similarity based on the calculated distance. This makes it possible to extend the period range of the learning data (combined data) when training a learning model, which is a time series prediction model, from the period range of the first data, thereby creating a learning model with higher quality.

[0094] (Appendix 5) The program described in Appendix 2, wherein the candidate extraction step (S103) includes a step of extracting first candidate data within a predetermined period range shorter than the third period range, and a step of extracting multiple candidate data included in the predetermined period range from the first candidate data by sequentially shifting the third data forward or backward by a predetermined period in the time direction. This makes it possible to obtain second data that is more similar to the first data from among multiple candidate data extracted by sequentially shifting each cycle period and included in a part of a third period that has a longer period range than the first data.

[0095] (Appendix 6) The program described in Appendix 5, wherein the processor executes a periodic input step (S103) of receiving input of a predetermined periodic period from a user, and a candidate extraction step (S103) of extracting multiple candidate data based on the periodic period received as input in the periodic input step. This allows the user to extract multiple candidate data that are sequentially shifted for each period (one day, one week, one month, etc.) according to a period appropriate for the type of the first data.

[0096] (Appendix 7) The program according to claim 5, wherein the candidate extracting step (S103) is a step of extracting a plurality of candidate data based on a periodic period specified based on the first data, without receiving an input of the periodic period from a user. This makes it possible to extract multiple candidate data that are sequentially shifted for each periodic period (one day, one week, one month, etc.) that is automatically identified based on the features of the first data, without receiving input from the user.

[0097] (Appendix 8) The program described in Appendix 2, wherein the candidate extraction step (S103) includes a step of excluding data that is earlier or later in the time direction than the first period range from the third data acquired in the third data acquisition step, and a step of extracting multiple candidate data from the excluded third data. This makes it possible to extend the period range of the learning data (combined data) from the period range of the first data when learning a learning model, which is a time-series prediction model, based on the second data, which does not include data that is earlier or later in the time direction than the period range of the first data, among the third data, and thus create a learning model of higher quality. For example, when constructing a time series prediction model for predicting events occurring after the first period (future) based on the first data, it is not preferable to use data from the third data that is after the period of the first data (future) in the time direction as learning data, in consideration of causal relationships. The same is true when constructing a time series prediction model for inferring events occurring before the first period (past) based on the first data.

[0098] (Appendix 9) The program described in Appendix 2, wherein the first data, the second data and the third data are time series data having a plurality of series, and the second data acquisition step (S104) is a step of acquiring the second data from the plurality of candidate data according to the similarity between the time series data of each series included in the first data and the time series data of each series included in the plurality of candidate data. As a result, when the first data, second data, and third data are time-series data having multiple series, the second data can be obtained according to the similarity of each series, allowing the creation of a higher quality learning model.

[0099] (Appendix 10) The program according to claim 9, wherein the second data acquiring step (S104) is a step of acquiring the second data from the plurality of candidate data according to the number of similar series between the first data and the plurality of candidate data. As a result, when the first data, second data, and third data are time-series data having multiple series, the candidate data with the largest number of similar series according to the similarity of each series can be obtained as the second data, making it possible to create a learning model with higher quality.

[0100] (Appendix 11) The program described in Appendix 1, wherein the processor executes an extended period receiving step (S105) of receiving from a user an extended period for the first data acquired in the first data acquisition step, and the combining step (S105) is a step of combining one or more second data at least either forward or backward in the time direction of the first data until the period range of the combined data reaches the extended period. This makes it possible to extend the period range of the learning data (combined data) when training a learning model, which is a time series prediction model, from the period range of the first data, thereby creating a learning model with higher quality.

[0101] (Appendix 12) The program according to claim 1, wherein the combining step (S105) is a step of creating combined data by combining one or more pieces of second data forward in a time direction of the first data. This makes it possible to extend the period range of the learning data (combined data) when training a learning model, which is a time series prediction model, from the period range of the first data, thereby creating a learning model with higher quality.

[0102] (Appendix 13) The program according to claim 1, wherein the combining step (S105) is a step of creating combined data by combining one or more second data items subsequent to the first data item in a time direction. This makes it possible to extend the period range of the learning data (combined data) when training a learning model, which is a time series prediction model, from the period range of the first data, thereby creating a learning model with higher quality.

[0103] (Appendix 14) A method implemented on a computer having a processor and a memory, the processor performing all the steps performed in the invention according to any one of claims 1 to 13. This makes it possible to extend the period range of the learning data (combined data) when training a learning model, which is a time series prediction model, from the period range of the first data, thereby creating a learning model with higher quality.

[0104] (Appendix 15) An information processing device comprising a control unit and a memory unit, wherein the control unit executes all of the steps executed in the invention according to any one of Supplementary Note 1 to Supplementary Note 13. This makes it possible to extend the period range of the learning data (combined data) when training a learning model, which is a time series prediction model, from the period range of the first data, thereby creating a learning model with higher quality.

[0105] (Appendix 16) A system comprising means for performing all the steps performed in any of the inventions according to appendix 1 to appendix 13. This makes it possible to extend the period range of the learning data (combined data) when training a learning model, which is a time series prediction model, from the period range of the first data, thereby creating a learning model with higher quality. [Explanation of symbols]

[0106] 1 system, 10 server, 101 storage unit, 104 control unit, 106 input device, 108 output device, 20 user terminal, 201 storage unit, 204 control unit, 206 input device, 208 output device

Claims

1. A program to be executed by a computer having a processor and a storage unit, The processor, A first data acquisition step of acquiring first data which is time series data; a second data acquiring step of acquiring second data having a time series transition similar to that of the first data from a plurality of time series data based on the first data acquired in the first data acquiring step; a combining step of generating combined data by combining the second data with at least one of a forward and a backward time direction of the first data; a learning step of learning a learning model based on the combined data created in the combining step; A program that executes.

2. The processor, a third data acquiring step of acquiring third data, which is time-series data having a third period range longer than a first period range of the first data acquired in the first data acquiring step; a candidate extraction step of extracting a plurality of candidate data which are a part of the third data acquired in the third data acquisition step and are a plurality of time-series data included in the third period range; Run The second data acquisition step is a step of acquiring the second data from the plurality of candidate data according to a similarity between the first data and the plurality of candidate data. The program according to claim 1.

3. The candidate extraction step is a step of extracting the plurality of candidate data having a period range substantially the same as the first period range. The program according to claim 2.

4. The second data acquisition step includes: Calculating at least one of a Manhattan distance, a Euclidean distance, a cosine similarity, and a correlation coefficient between the first data and the plurality of candidate data; acquiring the second data according to a similarity based on the calculated distance; Including, The program according to claim 2.

5. The candidate extraction step includes: extracting first candidate data within a predetermined period range shorter than the third period range; extracting a plurality of candidate data included in the predetermined period range from the first candidate data by sequentially shifting the third data forward or backward in a time direction by a predetermined period; Including, The program according to claim 2.

6. The processor, a cycle input step of receiving an input of the predetermined cycle period from a user; Run the candidate extraction step is a step of extracting the plurality of candidate data based on the periodic period inputted in the periodic input step; The program according to claim 5.

7. the candidate extraction step is a step of extracting the plurality of candidate data based on the cycle period specified based on the first data without receiving an input of a cycle period from a user; The program according to claim 5.

8. The candidate extraction step includes: excluding data that is before or after the first period range in a time direction from the third data acquired in the third data acquisition step; extracting a plurality of candidate data from the excluded third data; Including, The program according to claim 2.

9. the first data, the second data, and the third data are time series data having a plurality of series, the second data acquiring step is a step of acquiring the second data from the plurality of candidate data according to a similarity between each of the time series data of the plurality of candidate data and each of the time series data of the plurality of candidate data. The program according to claim 2.

10. The second data acquisition step is a step of acquiring the second data from the plurality of candidate data in accordance with the number of similar series between the first data and the plurality of candidate data. The program according to claim 9.

11. The processor, an extension period receiving step of receiving an extension period for the first data acquired in the first data acquiring step from a user; Run The combining step is a step of combining one or more of the second data with at least one of a front and a rear in a time direction of the first data until a period range of the combined data reaches the extended period. The program according to claim 1.

12. The combining step is a step of creating the combined data by combining one or more of the second data forward in a time direction of the first data. The program according to claim 1.

13. The combining step is a step of creating the combined data by combining one or more of the second data with the first data in a backward time direction. The program according to claim 1.

14. A method implemented on a computer having a processor and a memory, the processor performing all of the steps performed in any of the inventions according to claims 1 to 13.

15. 14. An information processing apparatus comprising a control unit and a storage unit, the control unit executing all steps executed in the invention according to any one of claims 1 to 13.

16. A system comprising means for performing all the steps performed in any one of claims 1 to 13.

Citation Information

Patent Citations

  • Abnormality detection method and system therefor

    JP2014149840A

  • Transferability determination device, transferability determination method and transferability determination program

    JP2021086241A

  • Learning model generation method and abnormality determination system

    JP2022135178A