Learning device, inference device, and training method

The learning device enhances event probability inference by extracting and labeling partial data from time-series data to generate a trained model, addressing the limitations of existing systems in predicting accidents or failures in industrial settings.

WO2026023099A1PCT designated stage Publication Date: 2026-01-29MITSUBISHI ELECTRIC CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/036917
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-24
Filing Date
2024-10-17
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Existing systems for predicting accidents or failures in refineries and chemical processing plants are unable to infer the probability of general events occurring based on time-series data.

Method used

A learning device that acquires time-series data, extracts partial data before and after specific events, and generates a trained model using logistic regression to infer the probability of future events by labeling first partial data as positive examples and second partial data as negative examples.

Benefits of technology

Enables accurate prediction of the likelihood of specific events by improving the inference model's accuracy through feature extraction and training data generation, allowing for better anticipation of future occurrences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024036917_29012026_PF_FP_ABST
    Figure JP2024036917_29012026_PF_FP_ABST
Patent Text Reader

Abstract

A learning device (100) comprises: a time-series data acquisition unit (11) that acquires time-series data in a specific period in which a plurality of specific events occurred; a partial data extraction unit (13) that divides the time-series data acquired by the time-series data acquisition unit (11) into pieces of partial data, and extracts therefrom pieces of first partial data respectively corresponding to first periods, which are periods immediately before the occurrences of the plurality of specific events, and pieces of second partial data respectively corresponding to second periods not overlapping the first periods of the plurality of specific events; and a training unit (18) that performs training on the basis of the pieces of first partial data and the pieces of second partial data by using the pieces of first partial data as positive examples and the pieces of second partial data as negative examples, and generates a trained model for inferring the probability that partial data of other time-series data corresponding to said time-series data is a positive example on the basis of an input of said partial data.
Need to check novelty before this filing date? Find Prior Art

Description

Learning device, inference device, and learning method

[0001] The present disclosure relates to a learning device, an inference device, and a learning method.

[0002] A computer system has been disclosed that predicts whether an accident or failure is likely to occur in a refinery or chemical processing plant based on time-series data acquired by multiple sensors (see, for example, Patent Document 1). This computer system identifies a precursor pattern of an accident or failure from the time-series data acquired by the multiple sensors, and predicts whether an accident or failure is likely to occur based on a dependency graph generated based on the time-series data and the precursor pattern.

[0003] International Publication No. 2018-009643

[0004] However, as mentioned above, the computer system described in Patent Document 1 predicts whether or not there is a possibility of an accident or failure occurring at a refinery or chemical processing plant based on time-series data acquired by multiple sensors, and has the problem of not being able to infer (predict) the probability of a general event occurring.

[0005] The present disclosure was made in recognition of the above-mentioned problem, and aims to provide a learning device, an inference device, and a learning method that can infer the probability of a specific event occurring.

[0006] The learning device according to the present disclosure is characterized by comprising: a time-series data acquisition unit that acquires time-series data for a specific period during which a plurality of specific events occurred; a partial data extraction unit that divides the time-series data acquired by the time-series data acquisition unit into a plurality of partial data and extracts, from the partial data, a plurality of first partial data corresponding to a first period that is a period immediately before the occurrence of each of the plurality of specific events and a plurality of second partial data corresponding to a plurality of second periods that do not overlap with each of the first periods of the plurality of specific events; and a learning unit that performs learning based on the plurality of first partial data and the plurality of second partial data, with each of the plurality of first partial data being a positive example and each of the plurality of second partial data being a negative example, and generates a trained model for inferring the probability that the partial data is a positive example based on input of partial data of other time-series data that corresponds to the time-series data.

[0007] The learning device according to the present disclosure can infer the probability of a particular event occurring.

[0008] 6A is a diagram showing a graph of first partial data extracted from the time series data by the learning device according to the embodiment 1, and FIG. 6B is a diagram showing a graph of second partial data extracted from the time series data by the learning device according to the embodiment 1. FIG. 6B is a diagram showing a graph of second partial data extracted from the time series data by the learning device according to the embodiment 1. FIG. 6C is a diagram showing a graph of the first partial data extracted from the time series data by the learning device according to the embodiment 1. FIG. 6D is a diagram showing a graph of the second partial data extracted from the time series data by the learning device according to the embodiment 1. FIG. 6E is a diagram showing a graph of the first partial data extracted from the time series data by the learning device according to the embodiment 1. FIG. 6F is a diagram showing a graph of the second partial data extracted from the time series data by the learning device according to the embodiment 1.

[0009] Embodiments of the present disclosure will be described in detail below with reference to the drawings. Embodiment 1. First, an inference system 1 according to embodiment 1 will be described with reference to FIG. 1. FIG. 1 is a block diagram showing a schematic configuration of the inference system 1 according to embodiment 1. As shown in FIG. 1, the inference system 1 according to embodiment 1 includes a learning device 100 and an inference device 500, which are connected wirelessly or via a wire so as to be able to communicate with each other. The learning device 100 is a device that generates a trained model for inferring the probability of a specific event occurring based on time-series data. The inference device 500 is a device that infers the probability of a specific event occurring using the trained model generated by the learning device 100. Note that it is sufficient for the inference device 500 to be able to acquire information from the learning device 100. The learning device 100 and the inference device 500 may be configured to be able to communicate information with each other via a device or communication line not shown.

[0010] The learning device 100 includes a time-series data acquisition unit 11 , a specific event information acquisition unit 12 , a partial data extraction unit 13 , a feature extraction unit 14 , a learning data generation unit 17 , and a learning unit 18 .

[0011] The time-series data acquisition unit 11 acquires time-series data for a specific period in which multiple specific events occurred. For example, the time-series data acquisition unit 11 acquires past time-series data that is empirically known to be correlated with the occurrence of multiple specific events during the specific period in which the multiple specific events occurred, or that can be empirically inferred to be correlated with the occurrence of multiple specific events. For example, although it is generally unclear what correlation exists between time-series data of a company's management indicators and important events of the company as specific events, it is empirically inferred that there is some correlation. The time-series data acquisition unit 11 acquires, for example, time-series data that can be inferred to be correlated with the occurrence of such specific events. Specifically, the time-series data acquisition unit 11 acquires time-series data indicating the trends in stock prices of a specific company. Note that the time-series data acquired by the time-series data acquisition unit 11 is not limited to this, and may be time-series data for a specific period in which multiple specific events occurred. For example, the time-series data acquisition unit 11 may be data indicating the time trends of economic indicators in a specific market, weather data indicating the time trends of observation results of specific weather information, or statistical data indicating the changes in social conditions in a specific society.

[0012] The specific event information acquisition unit 12 acquires information indicating the occurrence time of a specific event that occurred during a specific period indicated by the time series data acquired by the time series data acquisition unit 11. For example, if the time series data acquired by the time series data acquisition unit 11 is time series data indicating monthly changes in specific information, the specific event information acquisition unit 12 acquires information indicating the month in which the specific event occurred in the past, and if the time series data is time series data indicating time changes in the specific information, the specific event information acquisition unit 12 acquires information indicating the time in which the specific event occurred in the past. A specific event is an event whose occurrence probability is to be inferred by the inference device 500, and may be a sales activity such as a new product announcement at a company, an event such as the founding or bankruptcy of a company, or an event that occurs in society.

[0013] The partial data extraction unit 13 extracts partial data from the time series data acquired by the time series data acquisition unit 11 based on information indicating the occurrence time of specific events acquired by the specific event information acquisition unit 12. Specifically, the partial data extraction unit 13 divides the time series data acquired by the time series data acquisition unit 11 into multiple partial data, and extracts from the multiple partial data a multiple number of first partial data corresponding to a first period that is a period immediately before the occurrence of each of the multiple specific events, and a multiple number of second partial data corresponding to a multiple number of second periods that do not overlap with each of the first periods of the multiple specific events. Furthermore, the partial data extraction unit 13 extracts the multiple first partial data and the multiple second partial data such that each first period and each second period are the same length. Note that the multiple first periods and the multiple second periods may partially overlap each other. Furthermore, the first period may include the time point at which the specific event occurs, or may be long enough that the impact of the occurrence of the specific event and the time series data can be considered sufficiently large, for example, approximately one acquisition interval of the time series data, or the end of the first period may be before the time point at which the specific event occurs.

[0014] The feature extraction unit 14 extracts a feature from each of the plurality of first partial data and the plurality of second partial data extracted by the partial data extraction unit 13. Specifically, the feature extraction unit 14 extracts a feature indicating the degree of relevance of each of the plurality of first partial data and the plurality of second partial data with other partial data from the plurality of first partial data and the plurality of second partial data. In other words, the feature extraction unit 14 calculates, for each of the plurality of first partial data and the plurality of second partial data, a feature indicating the degree of relevance of each of the plurality of first partial data and the plurality of second partial data with other partial data. For example, the feature extraction unit 14 calculates, for each of the plurality of first partial data and the plurality of second partial data, a Discord score, which is a feature indicating the nearest distance to the plurality of other partial data, as a feature indicating the degree of relevance of each of the plurality of first partial data and the plurality of second partial data with other partial data.

[0015] The training data generation unit 17 generates training data for the learning device 100 to perform training, with the plurality of first partial data being positive examples and the plurality of second partial data being negative examples, based on the features extracted by the feature extraction unit 14. Specifically, for the plurality of features extracted from the plurality of first partial data and the plurality of second partial data by the feature extraction unit 14, the training data generation unit 17 assigns a positive example label to each of the feature quantities extracted from the plurality of first partial data and a negative example label to each of the feature quantities extracted from the plurality of second partial data, and generates training data in which the plurality of labeled feature quantities form a dataset.

[0016] The learning unit 18 performs learning based on the learning data generated by the learning data generation unit 17, and generates a trained model for inferring the probability that partial data is a positive example based on input of partial data of other time-series data corresponding to the time-series data acquired by the time-series data acquisition unit 11. For example, the learning unit 18 generates a logistic regression model for inferring the probability that input partial data is a positive example by logistic regression analysis based on the learning data generated by the learning data generation unit 17.

[0017] The inference device 500 includes a partial data acquisition unit 51, a feature extraction unit 52, an inference unit 53, and a storage unit 54 for storing information.

[0018] The partial data acquisition unit 51 acquires partial data used to perform inference using the trained model generated by the learning unit 18. For example, the partial data acquisition unit 51 acquires other time series data corresponding to the time series data acquired by the time series data acquisition unit 11, divides the acquired other time series data into multiple partial data, and extracts third partial data from the multiple partial data, which is partial data for a period to be inferred and has the same length as the first and second periods. For example, the partial data acquisition unit 51 acquires, as the partial data used to perform inference using the trained model, partial data representing the transition of the same data in a period after the specific period of time series data acquired by the learning device 100 to generate the training data. Furthermore, for example, the partial data acquisition unit 51 acquires, as the other time series data, time series data representing the transition of the same data in a period after the specific period of time series data acquired by the learning device 100 to generate the training data.

[0019] In addition, the partial data acquisition unit 51 is not limited to acquiring time series data showing the trend of the same data in a period after a specific period of the time series data acquired by the learning device 100 as partial data of other time series data corresponding to the time series data acquired by the learning device 100. The partial data acquisition unit 51 may acquire, as partial data of other time series data corresponding to the time series data acquired by the learning device 100, time series data that is related to changes in the time series data acquired by the learning device 100. For example, if the time series data acquired by the learning device 100 is time series data showing trends in representative values ​​of management indicators of multiple companies, the partial data acquisition unit 51 may be configured to acquire time series data showing trends in management indicators of a specific company in an industry that is common to the multiple companies; if the time series data acquired by the learning device 100 is time series data showing trends in management indicators of a specific company, the partial data acquisition unit 51 may be configured to acquire time series data showing trends in management indicators of other companies in an industry that is common to the specific company; or if the time series data acquired by the learning device 100 is time series data showing trends in economic indicators of a specific society, the partial data acquisition unit 51 may be configured to acquire time series data showing trends in economic indicators of other societies different from the specific society.

[0020] The feature extraction unit 52 extracts features from the partial data acquired by the partial data acquisition unit 51. Specifically, the feature extraction unit 52 extracts, from the partial data acquired by the partial data acquisition unit 51, features indicating the degree of relevance of each piece of partial data, which is made up of a plurality of first partial data and a plurality of second partial data extracted by the partial data extraction unit 13 of the learning device 100, with other partial data. In other words, the feature extraction unit 52 calculates, for each piece of partial data acquired by the partial data acquisition unit 51, a feature indicating the degree of relevance of each piece of partial data, which is made up of a plurality of first partial data and a plurality of second partial data extracted by the partial data extraction unit 13 of the learning device 100, with other partial data. For example, the feature extraction unit 52 calculates, for each piece of partial data acquired by the partial data acquisition unit 51, a Discord score, which is a feature indicating the nearest neighbor distance to the plurality of other partial data, as a feature indicating the degree of relevance of each piece of partial data with other partial data.

[0021] The inference unit 53 infers the probability that the partial data acquired by the partial data acquisition unit 51 is a positive example, using the trained model generated by the learning unit 18. In other words, the inference unit 53 infers the probability that partial data of other time-series data corresponding to the time-series data is a positive example, using the trained model generated by the learning unit 18. In other words, the inference unit 53 inputs the features extracted by the feature extraction unit 52 into the trained model generated by the learning unit 18, thereby acquiring, as an inference result, the probability that the partial data acquired by the partial data acquisition unit 51 is a positive example. The inference unit 53 outputs the inference result to another device, such as a display device (not shown) or another computer.

[0022] The storage unit 54 stores information used in various processes performed by the inference device 500 and information indicating the results of various processes performed by the inference device 500. Specifically, the storage unit 54 stores one or more of the time-series data, partial data, and feature amounts acquired by the learning device 100, the trained model generated by the learning device 100, the partial data acquired by the partial data acquisition unit 51, the feature amounts acquired by the feature extraction unit 52, and the results of inference by the inference unit 53. Each component of the inference device 500 refers to the information stored in the storage unit 54 when the inference device 500 performs processing.

[0023] Next, the hardware configuration of the learning device 100 will be described with reference to Figures 2 and 3. Figure 2 is a diagram showing an example of the hardware configuration of the learning device 100, and Figure 3 is a diagram showing an example of the hardware configuration of the learning device 100 that is different from that shown in Figure 2. For example, as shown in Figure 2, the learning device 100 is a computer having a processor 100a, a memory 100b, and an I / O port 100c, and is configured so that the processor 100a reads and executes a program stored in the memory 100b.

[0024] 3, the learning device 100 is a computer that has a processing circuit 100d, which is dedicated hardware, and an I / O port 100c, and executes a program. The processing circuit 100d is configured, for example, by a single circuit, a composite circuit, a programmed processor, a parallel programmed processor, an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or a combination thereof. Each function of the learning device 100 is realized by the processor 100a or the processing circuit 100d, which is dedicated hardware, executing a program. Note that the learning device 100 may also have hardware other than those described above, such as a hardware timer.

[0025] The hardware configuration of the inference device 500 is similar to that of the learning device 100, and therefore a description thereof will be omitted.

[0026] Next, the details of the processing performed by the learning device 100 and the processing performed by the inference device 500 will be described with reference to FIGS. 1 and 4 to 7. FIG. 4 is a flowchart showing an example of the processing performed by the learning device 100 according to embodiment 1. The processing performed by the learning device 100 shown in FIG. 4 is processing for generating a trained model that infers the probability of a specific event occurring based on time-series data. As shown in FIG. 4, when the learning device 100 starts the processing, it first acquires time-series data (step ST01). In this processing, the learning device 100 acquires the time-series data used to generate the training data using the time-series data acquisition unit 11.

[0027] 5 is a diagram showing a graph of time-series data for a specific period acquired by learning device 100 according to embodiment 1. As shown in Fig. 5, for example, time-series data J0 acquired by learning device 100 is data showing the transition of value k, which indicates specific information that changes over time t.

[0028] After performing the processing of step ST01, the learning device 100 acquires information indicating the occurrence times of specific events (step ST03). In this processing, the learning device 100 acquires information indicating the occurrence times of specific events that occurred during a specific period, which is the range of time t during which the time-series data J0 was acquired, using the specific event information acquisition unit 12. As shown in FIG. 5 , in this processing, for example, the learning device 100 acquires information indicating the occurrence times of specific events E1, E2, E3, and E4 that occurred during the specific period.

[0029] After performing the process of step ST03, the learning device 100 extracts first partial data and second partial data (step ST04). In this process, the learning device 100 extracts, for example, from the time-series data J0 acquired in the process of step ST01, multiple pieces of first partial data j1, j2, j3, and j4 corresponding to a first period p1, which is a period immediately before the occurrence of each of the multiple specific events E1 to E4, and multiple pieces of second partial data corresponding to multiple second periods that do not overlap with the first period p1 of each of the multiple specific events, based on information indicating the occurrence times of the specific events acquired in the process of step ST02, using the partial data acquisition unit 51.

[0030] 6A is a diagram showing a graph of first partial data j1 to j4 extracted from time-series data J0 by learning device 100 according to embodiment 1, and Fig. 6B is a diagram showing graphs of second partial data j5, j6, j7, j8, j9, j10, j11, j12, and j13 extracted from time-series data J0 by learning device 100 according to embodiment 1. As shown in Fig. 6A and Fig. 6B, in the process of step ST04, learning device 100 extracts, for example, from the graph of time-series data J0 acquired in the process of step ST01, a plurality of first partial data j1 to j4 represented by partial graphs (partial waveforms) of time-series data J0 and a plurality of second partial data j5 to j13 represented by partial graphs (partial waveforms) of time-series data J0.

[0031] After performing the process of step ST04, the learning device 100 extracts features from the first partial data j1 to j4 and the second partial data j5 to j13 (step ST06). In this process, the learning device 100 extracts features from each of the plurality of first partial data j1 to j4 and the plurality of second partial data j5 to j13 extracted in the process of step ST04 using the feature extraction unit 14.

[0032] As shown in Figures 6A and 6B, in the processing of step ST06, the learning device 100 extracts, for example, Discord scores of 4, 5, 8850, and 3024 as features from each of the multiple first partial data j1 to j4, and extracts Discord scores of 4, 4, 6, 5, 17963, 2410, 11, 24, and 8 as features from each of the multiple second partial data j5 to j13.

[0033] After performing the process of step ST06, the learning device 100 generates training data by labeling the extracted features (step ST11). In this process, the learning device 100 assigns labels of either positive examples or negative examples to the multiple features extracted in the process of step ST06, and generates training data with these multiple features as a data set.

[0034] After performing the process of step ST11, the learning device 100 generates a trained model (step ST13). In this process, the learning device 100 generates the trained model, which is a logistic regression model for inferring the probability that the first partial data is a positive example based on input of other partial data corresponding to the first partial data, for example, by logistic regression analysis based on the training data generated in the process of step ST11.

[0035] 7 is a flowchart showing an example of processing performed by the inference device 500 according to the first embodiment. The processing performed by the inference device 500 shown in FIG. 7 is processing for inferring the probability of a specific event occurring using a trained model generated by the learning device 100. As shown in FIG. 7, when the inference device 500 starts the processing, it first acquires a trained model (step ST21). In this processing, the inference device 500 acquires the trained model generated by the learning device 100 from the learning device 100.

[0036] After performing the process of step ST21, the inference device 500 acquires time-series data (step ST22). In this process, the inference device 500 acquires other time-series data corresponding to the time-series data acquired by the learning device 100 using the partial data acquisition unit 51.

[0037] After performing the process of step ST22, the inference device 500 extracts third partial data (step ST24). In this process, the inference device 500 divides the time-series data acquired in the process of step ST22 into multiple partial data, and acquires, from among the multiple partial data, the third partial data used for inference using the trained model generated by the learning device 100, using the partial data acquisition unit 51.

[0038] After performing the process of step ST24, the inference device 500 extracts features from the third partial data (step ST25). In this process, the inference device 500 extracts features from the third partial data acquired in the process of step ST24 using the feature extraction unit 52 to input the features to the trained model generated by the learning device 100.

[0039] After performing the process of step ST25, the inference device 500 inputs the extracted feature into the trained model (step ST26). In this process, the inference device 500 inputs the feature into the trained model generated by the learning device 100 in order to infer the probability that the partial data corresponding to the feature extracted in the process of step ST25 is a positive example.

[0040] After performing the process of step ST26, the inference device 500 acquires an inference result (step ST27). In this process, the inference device 500 acquires, as an inference result based on the trained model generated by the learning device 100, the probability that the third partial data acquired in the process of step ST24 is a positive example.

[0041] After performing the process of step ST27, the inference device 500 outputs the inference result (step ST28). In this process, the inference device 500 displays the inference result on a display device such as a liquid crystal display device that is connected to the inference device 500 so that information can be communicated.

[0042] As described above, the learning device 100 according to the first embodiment includes: a time-series data acquiring unit 11 that acquires time-series data during a specific period during which a plurality of specific events occurred; a partial data extracting unit 13 that divides the time-series data acquired by the time-series data acquiring unit 11 into a plurality of partial data and extracts, from the partial data, a plurality of first partial data corresponding to a first period that is a period immediately before the occurrence of each of the plurality of specific events and a plurality of second partial data corresponding to a plurality of second periods that do not overlap with each of the first periods of the plurality of specific events; and a learning unit 18 that performs learning based on the plurality of first partial data and the plurality of second partial data, with each of the plurality of first partial data being a positive example and each of the plurality of second partial data being a negative example, and generates a trained model for inferring the probability that the partial data is a positive example based on input of partial data of other time-series data that corresponds to the time-series data.

[0043] With this configuration, the learning device 100 can generate a trained model for inferring the probability that partial data of time-series data input to the trained model is a positive example, based on past time-series data and information indicating the time when a specific event occurred in the past. In other words, the learning device 100 can generate a trained model for inferring the probability that a specific event will occur immediately after partial data of time-series data input to the trained model, based on past time-series data and information indicating the time when a specific event occurred in the past.

[0044] For example, the learning device 100 can generate a trained model that infers the probability of a company in a specific industry going bankrupt in the near future based on time-series data indicating a market stock price index over a specific period in the past and information indicating, as a specific event, the timing of past company bankruptcies in the specific industry.Furthermore, for example, the learning device 100 can generate a trained model that infers the probability of a specific company releasing a new product in the near future based on time-series data indicating the stock price of a specific company over a specific period in the past and information indicating, as a specific event, the timing of the release of a new product by the specific company.

[0045] Furthermore, the learning device 100 according to the first embodiment includes a training data generation unit that generates training data for each of the plurality of first partial data and the plurality of second partial data based on feature quantities indicating the degree of association with other partial data included in the plurality of first partial data and the plurality of second partial data. For example, the learning device 100 according to the first embodiment is configured to generate training data for each of the plurality of first partial data and the plurality of second partial data based on the nearest neighbor distance to the other partial data as a feature quantity indicating the degree of association with the other partial data. With this configuration, the learning device 100 can generate a trained model in which the features of the partial data are more likely to be reflected in the inference result, thereby improving the accuracy of the inference result based on the trained model.

[0046] In the first embodiment, the learning device 100 generates training data for generating a trained model based on features extracted from the plurality of first partial data and the plurality of second partial data indicating the degree of association with other partial data, but is not limited to this. The learning device may be configured to perform training based on the plurality of first partial data and the plurality of second partial data, with each of the plurality of first partial data as a positive example and each of the plurality of second partial data as a negative example. For example, the learning device may be configured to label a feature indicating an average value (moving average) of each of the first partial data as a positive example and label a feature indicating an average value (moving average) of each of the second partial data as a negative example, and perform training based on training data in which the plurality of labeled features form a data set. Alternatively, the learning device may be configured to label a feature obtained by time-differentiating each of the first partial data as a positive example and label a feature obtained by time-differentiating each of the second partial data as a negative example, and perform training based on training data in which the plurality of labeled features form a data set. The learning device may also be configured to label a plurality of first partial data as positive examples, label a plurality of second partial data as negative examples, and perform learning based on learning data in which the labeled plurality of first partial data and the labeled plurality of second partial data form a data set.

[0047] Furthermore, in the first embodiment, the learning device 100 is configured to generate a trained model that is a logistic regression model, but is not limited to this. The learning device may be configured to generate a trained model that infers the probability that partial data of time-series data is a positive example. For example, the learning device may be configured to generate a trained model using an algorithm other than the logistic regression model, such as Naive Bayes, SVM (Support Vector Machine), Random Forest, or a deep learning algorithm, or may use a calibration method such as isotonic regression (see Non-Patent Document 1) or temperature scaling (see Non-Patent Document 2) in combination. Zadrozny, Bianca, et al. “Transforming classifier scores into accurate multiclass probability estimates.” Proceedings of the eighth ACM SIGKDD international conference on knowledge discovery and data mining. 2002. Hinton, George, et al. “Distilling the knowledge in a neural network.” arXiv preprint arXiv:1503.02531 (2015).

[0048] Second Embodiment Next, an inference system 2 according to a second embodiment will be described with reference to Figures 8 to 11. The inference system 2 according to the second embodiment differs in some configuration from the inference system according to the first embodiment, but other configurations are the same. The same components as those in the first embodiment are given the same names and symbols as those in the first embodiment, and descriptions thereof will be omitted.

[0049] FIG. 8 is a block diagram showing a schematic configuration of an inference system 2 according to embodiment 2. As shown in FIG. 8, the inference system 2 according to embodiment 2 includes a learning device 200 and an inference device 600, which are connected wirelessly or by wire so as to be able to communicate with each other. The learning device 200 is a device that generates a trained model for inferring the probability of a specific event occurring based on time-series data. The inference device 600 is a device that infers the probability of a specific event occurring using the trained model generated by the learning device 200.

[0050] The learning device 200 includes a time series data acquisition unit 11, a specific event information acquisition unit 12, a partial data extraction unit 13, a feature extraction unit 14, a relevance calculation unit 15, a time series data selection unit 16, a learning data generation unit 17, and a learning unit 18.

[0051] The relevance calculation unit 15 extracts a characteristic period from the specific period for which the time-series data was acquired by the time-series data acquisition unit 11, based on a feature indicating the relevance between multiple pieces of partial data included in the time-series data acquired by the time-series data acquisition unit 11, and calculates the relevance between the extracted characteristic period and multiple first periods. For example, the relevance calculation unit 15 extracts multiple pieces of partial data by sliding a window of a specific time width from the start point to the end point of the specific period, and extracts a feature indicating the relevance of each piece of partial data with other pieces of partial data. An example of a feature indicating the relevance of each piece of partial data with other pieces of partial data is the Discord score.

[0052] For example, the relevance calculation unit 15 extracts, as the characteristic period, a period in the time-series data that shows an outlier (outlier waveform) that is different from other partial data (partial waveforms). Specifically, the relevance calculation unit 15 extracts, as the characteristic period, a period in which the extracted feature quantity exceeds a preset threshold. More specifically, the relevance calculation unit 15 extracts, as one characteristic period, a period in which the extracted feature quantity continuously exceeds a preset threshold. Note that the relevance calculation unit 15 constitutes the period extraction unit in the first embodiment.

[0053] Furthermore, the relevance calculation unit 15 calculates the relevance between the extracted characteristic period and the plurality of first periods of the time series data acquired by the time series data acquisition unit 11. In other words, the relevance calculation unit 15 calculates a co-occurrence, which is the degree of co-occurrence between the extracted characteristic period and the plurality of first periods of the time series data acquired by the time series data acquisition unit 11. For example, the relevance calculation unit 15 calculates a recall Recall, which is an index indicating the reproducibility of the characteristic period in the first period, and indicates the proportion of periods that at least partially overlap with any of the plurality of first periods, as the relevance (co-occurrence) between the characteristic period and the first period. eve Calculate the recall rate Recall eve is calculated by the following formula (1): eve = (number of overlapping first periods and characteristic periods) / (number of first periods) (1)

[0054] Furthermore, for example, the relevance calculation unit 15 may use a discrimination ratio Precision, which is an index indicating the discrimination of the characteristic period in the first period, and indicates the proportion of periods among the plurality of characteristic periods that at least partially overlap with any of the first periods, as the relevance between the characteristic period and the first period. var Calculate the recall precision. var is calculated by the following formula (2): var = (number of overlaps between the first period and the characteristic period) / (number of characteristic periods) (2)

[0055] Furthermore, for example, the relevance calculation unit 15 may calculate the relevance between the characteristic period and the first period using the recall rate Recall eve and Precision var The harmonic mean D is calculated by the following formula (3): D=2×Recall eve ×Precision var / (Recall eve +Precision var ) ... (3)

[0056] When the time series data acquiring unit 11 acquires a plurality of pieces of time series data, the time series data selecting unit 16 selects one of the pieces of time series data from the acquired plurality of pieces of time series data based on the relevance calculated by the relevance calculating unit 15. For example, when the time series data acquiring unit 11 acquires first time series data and second time series data, the time series data selecting unit 16 selects, from these pieces of time series data, the time series data having the highest relevance calculated by the relevance calculating unit 15.

[0057] For example, the time-series data selection unit 16 selects, from the first time-series data and the second time-series data acquired by the time-series data acquisition unit 11, time-series data having high reproducibility calculated by the relevance calculation unit 15. Also, for example, the time-series data selection unit 16 selects, from the first time-series data and the second time-series data acquired by the time-series data acquisition unit 11, time-series data having high discriminability calculated by the relevance calculation unit 15. Also, for example, the time-series data selection unit 16 selects, from the first time-series data and the second time-series data acquired by the time-series data acquisition unit 11, time-series data having a high harmonic mean of the reproducibility and the discriminability calculated by the relevance calculation unit 15. When multiple pieces of time series data are acquired by the time series data acquisition unit 11, the time series data selection unit 16 may be configured to select, from the multiple pieces of time series data acquired, the time series data having the highest relevance calculated by the relevance calculation unit 15, or may be configured to select any one of the time series data based on the relevance calculated by the relevance calculation unit 15 and a condition other than the relevance calculated by the relevance calculation unit 15. The time series data selection unit 16 causes the learning data generation unit 17 to generate learning data based on the selected time series data.

[0058] The inference device 600 includes a partial data acquisition unit 61, a feature extraction unit 52, an inference unit 53, and a storage unit .

[0059] The partial data acquiring unit 61 acquires partial data used for performing inference using the trained model generated by the learning unit 18. For example, the partial data acquiring unit 61 acquires other time series data corresponding to the time series data selected by the time series data selecting unit 16, divides the acquired other time series data into a plurality of partial data, and extracts third partial data from the plurality of partial data, which is partial data for a period to be inferred and has the same length as the first period and the second period. Other features of the partial data acquiring unit 61 are similar to those of the partial data acquiring unit 51 according to the first embodiment, and therefore description thereof will be omitted.

[0060] The hardware configurations of the learning device 200 and the inference device 600 are similar to that of the learning device 100, and therefore will not be described here.

[0061] Next, the processing performed by the learning device 200 and the processing performed by the inference device 600 will be described in detail with reference to Figures 8 to 11. Figure 9 is a flowchart showing an example of processing performed by the learning device 200 according to embodiment 2. The processing performed by the learning device 200 shown in Figure 9 is processing for generating a trained model that infers the probability of a specific event occurring based on time-series data. Note that part of the processing performed by the learning device 200 according to embodiment 2 is similar to the processing performed by the learning device 100 according to embodiment 1, and therefore, a description of processing similar to that of embodiment 1 will be omitted.

[0062] 9 , when the learning device 200 starts the process, it first acquires first time series data and second time series data (step ST02). In this process, the learning device 200 acquires the first time series data and the second time series data as time series data used to generate learning data using the time series data acquisition unit 11. Note that in this process, the learning device 200 may be configured to acquire three or more pieces of time series data including the first time series data and the second time series data using the time series data acquisition unit 11.

[0063] After performing the process of step ST02, the learning device 200 acquires information indicating the occurrence time of the specific event (step ST03).

[0064] After performing the process of step ST03, the learning device 200 extracts first partial data and second partial data from the first time-series data and the second time-series data (step ST05). In this process, the learning device 200 extracts, from each of the first time-series data and the second time-series data acquired in the process of step ST02, a plurality of first partial data corresponding to a first period that is a period immediately before the occurrence of each of the plurality of specific events, and a plurality of second partial data corresponding to a plurality of second periods that do not overlap with each of the first periods of each of the plurality of specific events, using the partial data extraction unit 13.

[0065] After performing the process of step ST05, the learning device 200 extracts features from the first partial data and the second partial data (step ST07). In this process, the learning device 200 extracts features from each of the plurality of first partial data and the plurality of second partial data extracted in the process of step ST05 using the feature extraction unit 14.

[0066] 10 is a diagram showing graphs of feature quantities of the first time series data J1 and the second time series data J2 over a specific period acquired by the learning device 200 according to embodiment 1. For example, in the graph showing the feature quantities of the first time series data J1, portions S11, S12, S13, S14, S15, S16, and S17 show that the feature quantities exceed a threshold value th1 that is preset for each piece of time series data. Also, in the graph showing the feature quantities of the second time series data J2, portions S21, S22, S23, S24, and S25 show that the feature quantities exceed a threshold value th2 that is preset for each piece of time series data.

[0067] After performing the process of step ST07, the learning device 200 calculates a recall rate indicating the reproducibility of the characteristic periods for each of the first time-series data and the second time-series data (step ST08). In this process, the learning device 200 calculates the recall rate, which is the proportion of periods that at least partially overlap with the seven characteristic periods corresponding to S11 to S17 among the eight first periods corresponding to the specific events E11, E12, E13, E14, E15, E16, E17, and E18 in the graph of FIG. 10, to be 6 / 8.

[0068] After performing the process of step ST08, the learning device 200 calculates the distinctiveness of the characteristic periods for each of the first time-series data and the second time-series data (step ST09). In this process, the learning device 200 calculates a distinctiveness ratio, which is the proportion of periods that at least partially overlap with the eight first periods, among the seven characteristic periods corresponding to S11 to S17 in the graph of FIG. 10, to be 1 / 7.

[0069] After performing the process of step ST09, the learning device 200 calculates the degree of association between the specific event and the characteristic period (step ST10). In this process, the learning device 200 calculates the degree of association between the specific event and the characteristic period, for example, by calculating the harmonic mean of the recall rate calculated in the process of step ST08 and the discrimination rate calculated in the process of step ST09.

[0070] After performing the process of step ST10, the learning device 200 generates learning data by labeling the feature quantities extracted from the time-series data having a high degree of association between the specific event and the characteristic period from among the first time-series data and the second time-series data (step ST12). Note that, as described above, the learning device 200 may be configured to select one of the plurality of time-series data using the recall rate calculated in the process of step ST08 and the discrimination rate calculated in the process of step ST09 as the degree of association between the specific event and the characteristic period, or may be configured to select one of the plurality of time-series data using the arithmetic mean (arithmetic mean) of the recall rate calculated in the process of step ST08 and the discrimination rate calculated in the process of step ST09 as the degree of association between the specific event and the characteristic period, or may be configured to select one of the plurality of time-series data based on another degree of association.

[0071] After performing the process of step ST12, the learning device 200 generates a trained model (step ST13). The process of generating a trained model based on training data by the learning device 200 is the same as that of the learning device 200 according to the first embodiment.

[0072] Fig. 11 is a flowchart showing an example of processing performed by inference device 600 according to embodiment 2. The processing performed by inference device 600 shown in Fig. 11 is processing for inferring the probability of a specific event occurring using a trained model generated by learning device 200. As shown in Fig. 11, when inference device 600 starts processing, it first acquires a trained model (step ST21).

[0073] After performing the process of step ST21, the inference device 600 acquires time series data corresponding to the time series data having a high degree of association between the specific event and the characteristic period from among the first time series data and the second time series data (step ST23). In this process, the inference device 600 acquires time series data corresponding to the time series data selected based on the degree of association between the specific event and the characteristic period from among the multiple time series data acquired by the learning device 200, in order to infer the probability of the specific event occurring using the trained model generated by the learning device 200.

[0074] After performing the process of step ST23, the inference device 600 extracts third partial data (step ST24). In this process, the inference device 600 divides the time-series data acquired in the process of step ST23 into multiple partial data, and the partial data acquisition unit 61 acquires the third partial data from among the multiple partial data to be used for inference using the trained model generated by the learning device 100.

[0075] After performing the processing of step ST24, the inference device 600 extracts features from the third partial data (step ST25). After performing the processing of step ST25, the inference device 600 inputs the extracted features into the trained model (step ST26). After performing the processing of step ST26, the inference device 600 acquires an inference result (step ST27). After performing the processing of step ST27, the inference device 600 outputs the inference result (step ST28).

[0076] As described above, the learning device 200 according to the second embodiment includes a relevance calculation unit 15 that extracts characteristic feature periods within a specific period based on feature quantities indicating the relevance between multiple partial data included in time series data, and calculates the relevance between the extracted feature periods and multiple first periods. The learning device 200 is configured to perform learning based on partial data of time series data with a high relevance calculated by the relevance calculation unit 15, from among the first time series data and second time series data acquired by the time series data acquisition unit.

[0077] With this configuration, when the learning device 200 is able to acquire multiple time-series data for inferring the probability of a specific event occurring, it can select time-series data that has a high correlation between a characteristic period of the time-series data and the occurrence of the specific event, and generate a trained model based on the selected time-series data. This can improve the accuracy of inferring the probability of a specific event occurring using the trained model.

[0078] Furthermore, the learning device 200 according to the second embodiment is configured to calculate the degree of relevance between a characteristic period extracted by the relevance calculation unit 15 and a plurality of first periods, based on the proportion of overlap between the plurality of first periods and the characteristic periods extracted by the relevance calculation unit 15. Furthermore, the learning device 200 according to the second embodiment is configured to calculate the degree of relevance between a characteristic period extracted by the relevance calculation unit 15 and a plurality of first periods, based on the proportion of overlap between the plurality of first periods and the characteristic periods extracted by the relevance calculation unit 15. With this configuration, the learning device 200 can select time-series data suitable for generating a trained model from a plurality of time-series data, based on the reproducibility or distinctiveness of characteristic periods in the time-series data.

[0079] In the second embodiment, the learning device 200 is configured to perform learning based on partial data of time-series data that has the highest degree of association between a characteristic period and a first period among the acquired multiple time-series data, but not based on partial data of other time-series data. The learning device may be configured to perform learning based on partial data of time-series data selected based on the degree of association between a characteristic period and a first period among the acquired multiple time-series data. For example, the learning device may be configured to perform learning based on partial data of time-series data other than the time-series data that has the lowest degree of association between a characteristic period and a first period among the acquired multiple time-series data. Alternatively, the learning device may be configured to perform learning based on partial data of some time-series data selected based on the degree of association between a characteristic period and a first period among the acquired multiple time-series data, or to weight differently time-series data that has a relatively high degree of association between a characteristic period and a first period and time-series data that has a relatively low degree of association among the acquired multiple time-series data, and perform learning based on the partial data of each time-series data.

[0080] The learning device may have some or all of the configuration of the inference device, or the inference device may have some or all of the configuration of the learning device, or the learning device or some of the configuration of the inference device may be provided in another device that is communicatively connected to the learning device and the inference device. Also, some of the configuration of the learning device may have some of the functions of the inference device, or some of the configuration of the inference device may have some of the functions of the learning device.

[0081] In addition, the present disclosure allows for free combination of the respective embodiments, modification of any of the components of the respective embodiments, or omission of any of the components of the respective embodiments.

[0082] An information processing device according to the present disclosure can be used to infer the probability of a specific event that will affect a company's performance occurring based on time-series data such as economic indicators, for example.

[0083] Various aspects of the present disclosure are summarized below as appendices.

[0084] a learning unit configured to perform learning based on the plurality of first partial data and the plurality of second partial data, with each of the plurality of first partial data being a positive example and each of the plurality of second partial data being a negative example, and to generate a trained model for inferring a probability that a partial data piece is a positive example based on input of partial data of other time series data corresponding to the time series data. (Supplementary Note 1) A learning device comprising: a time-series data acquisition unit that acquires time-series data for a specific period during which a plurality of specific events occurred; a partial data extraction unit that divides the time-series data acquired by the time-series data acquisition unit into a plurality of partial data and extracts from the plurality of partial data: a plurality of first partial data corresponding to a first period that is a period immediately before the occurrence of each of the plurality of specific events; and a learning data generation unit that generates the training data for each of the plurality of first partial data and the plurality of second partial data based on a feature that indicates a degree of relevance between the plurality of first partial data and other partial data included in the plurality of second partial data. (Supplementary Note 3) The learning device according to Supplementary Note 1 or 2, wherein the learning data generation unit generates the learning data for each of the plurality of first partial data and the plurality of second partial data based on a nearest neighbor distance to other partial data as a feature indicating a degree of relevance with other partial data. (Supplementary Note 4) The learning device according to any one of Supplementary Notes 1 to 3, further comprising: a period extraction unit that extracts a characteristic period in the specific period based on a feature indicating a degree of relevance between the plurality of partial data included in the time series data, and an association calculation unit that calculates an association between the characteristic period extracted by the period extraction unit and the plurality of first periods, wherein the learning unit performs learning based on partial data of time series data with a high degree of relevance calculated by the association calculation unit, of the first time series data and the second time series data acquired by the time series data acquisition unit.(Supplementary Note 5) The learning device according to any one of Supplements 1 to 4, wherein the learning unit does not perform learning based on partial data of time series data having a low degree of relevance calculated by the relevance calculation unit, among the first time series data and the second time series data acquired by the time series data acquisition unit. (Supplementary Note 6) The learning device according to any one of Supplements 1 to 5, wherein the relevance calculation unit calculates the degree of relevance between a characteristic period extracted by the period extraction unit and the multiple first periods, based on a proportion of the multiple first periods that overlap with the multiple characteristic periods extracted by the period extraction unit. (Supplementary Note 7) The learning device according to any one of Supplements 1 to 6, wherein the relevance calculation unit calculates the degree of relevance between a characteristic period extracted by the period extraction unit and the multiple first periods, based on a proportion of the multiple first periods that overlap with the multiple first periods, among the multiple characteristic periods extracted by the period extraction unit. (Supplementary Note 8) An inference device comprising: a memory unit that stores a trained model generated by learning based on a plurality of first partial data and a plurality of second partial data, the trained model being extracted from time series data during a specific period in which a plurality of specific events occurred, and corresponding to a first period that is a period immediately before the occurrence of each of the plurality of specific events, as a positive example, and a plurality of second partial data corresponding to a plurality of second periods that do not overlap with each of the periods of the plurality of specific events, and the trained model is used to infer a probability that partial data of other time series data corresponding to the input time series data is a positive example; and an inference unit that infers a probability that partial data of other time series data corresponding to the time series data is a positive example using the trained model.(Supplementary Note 9) A learning method performed by an apparatus including a time-series data acquiring unit, a partial data extracting unit, and a learning unit, comprising: a step in which the time-series data acquiring unit acquires time-series data for a specific period in which a plurality of specific events occurred; a step in which the partial data extracting unit divides the time-series data acquired by the time-series data acquiring unit into a plurality of partial data, and extracts, from the partial data, a plurality of first partial data corresponding to a first period that is a period immediately before the occurrence of each of the plurality of specific events, and a plurality of second partial data corresponding to a plurality of second periods that do not overlap with each of the first periods of the plurality of specific events; and a step in which the learning unit performs learning based on the plurality of first partial data and the plurality of second partial data, with each of the plurality of first partial data as a positive example and each of the plurality of second partial data as a negative example, and generates a trained model for inferring the probability that the partial data is a positive example based on input of partial data of other time-series data that corresponds to the time-series data.

[0085] 1 Inference system, 2 Inference system, 11 Time series data acquisition unit, 12 Specific event information acquisition unit, 13 Partial data extraction unit, 14 Feature extraction unit, 15 Relevance calculation unit (period extraction unit), 16 Time series data selection unit, 17 Learning data generation unit, 18 Learning unit, 51 Partial data acquisition unit, 52 Feature extraction unit, 53 Inference unit, 54 Memory unit, 61 Partial data acquisition unit, 100 Learning device, 100a Processor, 100b Memory, 100c I / O port, 100d Processing circuit, 200 Learning device, 500 Inference device, 600 Inference device, D Harmonic mean, E1 Specific event, E11 Specific event, E12 Specific event, E13 Specific event, E14 Specific event, E15 Specific event, E16 Specific event, E17 Specific event, E18 Specific event, E2 Specific event, E3 specific event, E4 specific event, J0 time series data, J1 first time series data, J2 second time series data, j1 first partial data, j10 second partial data, j11 second partial data, j12 second partial data, j13 second partial data, j2 first partial data, j3 first partial data, j4 first partial data, j5 second partial data, j6 second partial data, j7 second partial data, j8 second partial data, j9 second partial data, p1 first period, th1 threshold, th2 threshold.

Claims

1. A learning device comprising: a time series data acquisition unit that acquires time series data during a specific period in which a plurality of specific events occurred; a partial data extraction unit that divides the time series data acquired by the time series data acquisition unit into a plurality of partial data and extracts from the partial data a plurality of first partial data corresponding to a first period that is a period immediately before the occurrence of each of the plurality of specific events, and a plurality of second partial data corresponding to a plurality of second periods that do not overlap with each of the first periods of the plurality of specific events; and a learning unit that performs learning based on the plurality of first partial data and the plurality of second partial data, using each of the plurality of first partial data as a positive example and each of the plurality of second partial data as a negative example, and generates a trained model for inferring the probability that the partial data is a positive example based on input of partial data of other time series data that corresponds to the time series data.

2. The learning device described in claim 1, characterized in that it is provided with a learning data generation unit that generates learning data for each of the plurality of first partial data and the plurality of second partial data based on features that indicate the degree of relevance with other partial data contained in the plurality of first partial data and the plurality of second partial data.

3. The learning device described in claim 2, characterized in that the learning data generation unit generates the learning data for each of the plurality of first partial data and the plurality of second partial data based on the nearest neighbor distance to other partial data as a feature indicating the degree of relevance with other partial data.

4. A learning device as described in any one of claims 1 to 3, characterized in that it comprises: a period extraction unit that extracts a characteristic feature period in the specific period based on a feature indicating the degree of relevance between multiple partial data included in the time series data; and a relevance calculation unit that calculates the degree of relevance between the characteristic period extracted by the period extraction unit and the multiple first periods, and the learning unit performs learning based on partial data of time series data that has a high degree of relevance calculated by the relevance calculation unit, out of the first time series data and second time series data acquired by the time series data acquisition unit.

5. The learning device described in claim 4, characterized in that the learning unit does not perform learning based on partial data of time series data that has a low relevance calculated by the relevance calculation unit, among the first time series data and second time series data acquired by the time series data acquisition unit.

6. A learning device as described in claim 4 or 5, characterized in that the relevance calculation unit calculates the relevance between the characteristic period extracted by the period extraction unit and the multiple first periods based on the proportion of the multiple first periods that overlap with the multiple characteristic periods extracted by the period extraction unit.

7. A learning device described in any one of claims 4 to 6, characterized in that the relevance calculation unit calculates the relevance between the characteristic period extracted by the period extraction unit and the multiple first periods based on the proportion of overlap between the characteristic period extracted by the period extraction unit and the multiple first periods.

8. An inference device comprising: a memory unit that stores a trained model generated by learning based on multiple first partial data and multiple second partial data, the trained model being extracted from time series data during a specific period in which multiple specific events occurred, and corresponding to a first period that is a period immediately before the occurrence of each of the multiple specific events, as a positive example, and multiple second partial data corresponding to a second period that does not overlap with each of the periods of the multiple specific events, and the trained model is used to infer the probability that partial data of other time series data corresponding to the input time series data is a positive example; and an inference unit that infers the probability that partial data of other time series data corresponding to the time series data is a positive example using the trained model.

9. A learning method performed by an apparatus including a time-series data acquisition unit, a partial data extraction unit, and a learning unit, comprising: a step in which the time-series data acquisition unit acquires time-series data for a specific period in which a plurality of specific events occurred; a step in which the partial data extraction unit divides the time-series data acquired by the time-series data acquisition unit into a plurality of partial data, and extracts from the partial data: a plurality of first partial data corresponding to a first period that is a period immediately before the occurrence of each of the plurality of specific events, and a plurality of second partial data corresponding to a plurality of second periods that do not overlap with each of the first periods of the plurality of specific events; and a step in which the learning unit performs learning based on the plurality of first partial data and the plurality of second partial data, using each of the plurality of first partial data as a positive example and each of the plurality of second partial data as a negative example, and generates a trained model for inferring the probability that the partial data is a positive example based on input of partial data of other time-series data that corresponds to the time-series data.

Citation Information

Patent Citations

  • Fault prediction system, fault prediction device and program

    JP2015174256A