Data extraction device, learning model construction device, data extraction method, learning model construction method, and program

By employing similarity-based data reduction and classification methods, the time-consuming process of learning model construction is expedited without compromising accuracy, utilizing normalized cross-correlation and frequency distribution to streamline the deep learning process.

JP7702799B2Active Publication Date: 2025-07-04MITSUBISHI HEAVY IND LTD
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
JP2021062980
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-04-01
Publication Date
2025-07-04
Estimated Expiration
2041-04-01

AI Technical Summary

Technical Problem

The construction of learning models, particularly using deep learning, requires significant time due to the large amount of learning data and the need to account for temporal differences in time-series data.

Method used

A data extraction method that reduces learning data by calculating similarity based on features excluding temporal differences, using normalized cross-correlation and frequency distribution, and classifying data into groups to delete redundant data, followed by constructing a learning model with the reduced data.

Benefits of technology

This approach significantly reduces the time required for learning model construction while maintaining accuracy by minimizing redundant data and focusing on essential features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007702799000001
    Figure 0007702799000001
  • Figure 0007702799000002
    Figure 0007702799000002
  • Figure 0007702799000003
    Figure 0007702799000003
Patent Text Reader

Abstract

To provide a data extraction method for reducing learning data.SOLUTION: A data extraction method which extracts a part of learning data from a plurality of learning data extracted for each time window from time-series data comprises the steps of: acquiring the plurality of learning data; excluding a temporal difference before and after the time window to calculate degree of similarity between the learning data on the basis of a feature of the learning data; and separating the learning data having high degree of similarity to delete a part thereof.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a data extraction device, a learning model construction device, a data extraction method, a learning model construction method, and a program.

Background Art

[0002] In recent years, the use of deep learning has been on the rise. For example, Patent Document 1 discloses a method for improving detection accuracy in target detection within an image using deep learning. Also, for example, when dealing with a multi-class classification problem of time-series data, high accuracy may be obtained by using deep learning. For example, the application of CNN (CNN: Convolutional Neural Network) is expected. Learning and evaluation of CNN require a great deal of time. For example, when constructing a discrimination model for discriminating the operating state of a plant from the process data of the plant, if a discrimination model is to be constructed for each plant, an enormous amount of time is required for learning and evaluation.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] There is a need for a method to reduce the time required for constructing a learning model.

[0005] The present disclosure provides a data extraction device, a learning model construction device, a data extraction method, a learning model construction method, and a program that can solve the above problems.

Means for Solving the Problems

[0006] The data extraction device of the present disclosure is a data extraction device that extracts some learning data from a plurality of learning data cut out for each time window from time series data, and includes a data acquisition unit that acquires the plurality of learning data, and a similarity calculation unit that calculates the similarity between the learning data based on the features of the learning data excluding the temporal differences before and after in the time window, and a data reduction unit that separates the learning data with high similarity and deletes a part of them. Further, the similarity calculation unit may calculate the similarity based on the magnitude per unit time of the overlapping area of the regions related to the two learning data with respect to the area sandwiched between the waveform data indicated by each learning data when the two learning data are shifted in the time axis direction. Further, the similarity calculation unit may calculate the similarity between the two shapes based on the shape of each frequency distribution obtained by performing frequency analysis on the two learning data. Further, it further includes a classification unit that classifies the plurality of learning data into the same group based on a predetermined evaluation criterion, and the similarity calculation unit may calculate the similarity between the learning data belonging to the same group. 。

[0007] The learning model construction device of the present disclosure includes the above data extraction device and a learning unit that learns the learning data to construct a learning model. The learning unit learns the learning data extracted by the data extraction device to construct the learning model, and then adjusts the parameters of the learning model using the learning data before extraction.

[0008] The data extraction method of the present disclosure is a data extraction method for extracting some learning data from a plurality of learning data cut out for each time window from time-series data, and includes steps of: obtaining a plurality of the learning data; calculating a similarity between the learning data based on features of the learning data excluding temporal differences before and after in the time window; separating the learning data with high similarity from each other and deleting a part of them. Further, in the step of calculating the similarity, when two of the learning data are shifted in the time axis direction, the similarity may be calculated based on the magnitude per unit time of the overlapping area of the regions sandwiched by the waveform data indicated by each learning data and the time axis. Also, in the step of calculating the similarity, the similarity between the two shapes may be calculated based on the shape of each frequency distribution obtained by frequency analyzing two of the learning data. Further, the method further includes a step of classifying the plurality of the learning data into the same group based on a predetermined evaluation scale, and in the step of calculating the similarity, the similarity between the learning data belonging to the same group may be calculated. 。

[0009] The learning model construction method of the present disclosure constructs a learning model by learning the learning data extracted by the above data extraction method, and then adjusts the parameters of the learning model using the learning data before extraction.

[0010] The program of the present disclosure is a data extraction process for a computer to extract some learning data from a plurality of learning data cut out for each time window from time-series data, including steps of acquiring the plurality of learning data, calculating the similarity between the learning data based on the features of the learning data excluding the temporal differences before and after in the time window, and separating and deleting a part of the learning data with high similarity. Further, in the step of calculating the similarity, when two pieces of the learning data are shifted in the time axis direction, the similarity may be calculated based on the magnitude per unit time of the overlapping area between the regions sandwiched by the waveform data indicated by each learning data and the time axis. Also, in the step of calculating the similarity, the similarity between the two shapes may be calculated based on the shape of each frequency distribution obtained by performing frequency analysis on the two pieces of the learning data. Further, the data extraction process executed by the program further includes a step of classifying the plurality of learning data into the same group based on a predetermined evaluation criterion for those that are similar, and in the step of calculating the similarity, the similarity between the learning data belonging to the same group may be calculated. 。

[0011] The program of the present disclosure causes a computer to execute a process of learning the learning data extracted by the above data extraction process to construct a learning model, and then adjusting the parameters of the learning model using the learning data before extraction.

Advantages of the Invention

[0012] According to the above data extraction device, learning model construction device, data extraction method, learning model construction method, and program, by reducing the learning data, the time required for constructing the learning model can be reduced.

Brief Description of the Drawings

[0013]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5A

Figure 5B

Figure 5C

Figure 5D

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

MODE FOR CARRYING OUT THE INVENTION

[0014] <First Embodiment> Hereinafter, a data extraction method and a learning model construction method according to the first embodiment of the present disclosure will be described with reference to FIGS. 1 to 7. (Configuration) FIG. 1 is a block diagram showing an example of a learning model construction device according to the embodiment. As shown in the figure, the learning model construction device 10 includes a data acquisition unit 11, a data extraction unit 12, a learning unit 13, an input unit 14, an output unit 15, and a storage unit 16. The data acquisition unit 11 acquires learning data used for constructing a learning model. The learning data is, for example, data cut out for each time window of a predetermined time width from time series data such as temperature, pressure, and flow rate collected from a device or plant to be monitored. Here, refer to FIG. 2. An example of the process (random sampling) of cutting out learning data from time series data is shown in the upper diagram 201 of FIG. 2. Variables A to C are time series data of various parameters such as temperature, pressure, and flow rate. In random sampling, learning data is cut out one by one in a time window having a common start time and end time from the time series data of each parameter, and the time window is shifted arbitrarily to cut out a plurality of learning data. The data acquisition unit 11 acquires data cut out in a common time window for each of variables A to C. These data are used as learning data for constructing a learning model. Here, variables A to C are used as examples for explanation, but the number of variables is not limited to three, and may be two or less or four or more.

[0015] The data extraction unit 12 extracts data from the training data acquired by the data acquisition unit 11 so that the learning accuracy does not decrease. Here, refer to the lower diagram 202 in FIG. 2. The data extraction unit 12 extracts only representative data (for example, W1) that does not redundantly include data (for example, W1 and W4) that is determined to have the same characteristics in time series among the training data cut out for each time frame W1 to W9 as the training data. The data extraction unit 12 includes a similarity calculation unit 121 and a data reduction unit 122. The similarity calculation unit 121 calculates the similarity of data based on (1) normalized cross-correlation, (2) frequency distribution, etc. The method for calculating the similarity will be described later with reference to FIGS. 3 to 4. The data reduction unit 122 reduces data with overlapping features in time series from the training data based on the similarity between the data. The method for reducing data from the training data will be described later with reference to FIGS. 5A to 5D. Here, data with overlapping features in time series means not only cases where the data cut out for each time frame are similar as they are, but also cases where the data cut out for each time frame are shifted in the time axis direction (not shifting the time frame in the time axis direction, but shifting the data after cutting out with a common time frame).) and the features are similar. When constructing a learning model in which the time series information (information before and after time) of the training data is excluded and learned, such as deep learning like CNN, even if data with overlapping features in time series is reduced, if a small amount of training data having such features is left, it is considered that a learning model corresponding to such features can be constructed. By excluding data with overlapping features in time series from the training data and narrowing down to a small number of training data having the features necessary for constructing a learning model from the entire training data and performing learning, the learning time can be shortened.

[0016] The learning unit 13 learns the training data by deep learning and constructs a learning model. The input unit 14 is configured using input devices such as a keyboard, a mouse, a touch panel, and buttons. The input unit 14 receives information input using the input device and outputs the information to the data extraction unit 12, the learning unit 13, etc. The output unit 15 outputs various types of data to an electronic file or a display device. For example, the output unit 15 may output the learning data extracted by the data extraction unit 12 as an electronic file. The storage unit 16 stores data necessary for extraction of learning data and construction of a learning model.

[0017] [Calculation of similarity] The similarity calculation unit 121 calculates the similarity between learning data using the following normalized cross-correlation or frequency component distribution.

[0018] (1) Similarity based on normalized cross-correlation Next, a method for calculating similarity based on normalized cross-correlation will be described with reference to FIGS. 2 and 3. The waveform data L1 and waveform data L2 in FIG. 3 are data obtained by normalizing the data of variable A in time frame W1 and the data of variable A in time frame W4, respectively, cut out from time-series data so that the average value of the data values is 0 and the standard deviation is 1. The similarity calculation unit 121 calculates, for example, the area P1 of the overlapping portion of the waveforms of the waveform data L2 while shifting the waveform data L2 in the time axis direction with respect to the waveform data L1. Then, the similarity calculation unit 121 divides the area P1 of the overlapping portion by the time T1 of the overlapping portion to calculate the unit area of the overlapping portion. Similarly to variable A, the similarity calculation unit 121 calculates waveform data obtained by normalizing the data of variable B and variable C in time frame W1 and the data in time frame W4, respectively, and calculates the unit area of the overlapping portion of the waveform data calculated for time frame W1 and the waveform data calculated for time frame W4 when time frame W4 is shifted in the time direction by the same amount as variable A. Let the unit area for variable B at this time be P2 and the unit area for variable C be P3. The similarity calculation unit 121 calculates the average value of the unit areas P1 to P3 while adjusting the amount of shift in the time direction of the data in time frame W4, and searches for the amount of shift in time frame W4 that maximizes the average value. Then, the similarity calculation unit 121 sets the average value when the average value of the unit areas P1 to P3 is maximized as the similarity between the data in time frame W1 and the data in time frame W4. By such calculation, even for waveforms with a phase difference, if there is a period with a similar trend, the similarity value will increase. The similarity calculation unit 121 calculates the similarity between time frame W1 and time frames W2 to W9 in the same way, calculates the similarity between time frame W2 and time frames W3 to W9, and calculates the similarity for all patterns of time frame combinations in the same way. When all the time-series values of one piece of learning data are the same value, normalization cannot be performed, so the similarity of the calculation target variable is set to 1.

[0019] (2) Similarity Based on Frequency Component Distribution Next, with reference to FIG. 4, a method for calculating similarity based on the distribution of frequency components will be described. For example, the similarity calculation unit 121 performs frequency analysis by fast Fourier transform on each of the data of variable A cut out in time frame W1 that has been normalized and the data of variable A cut out in time frame W4 that has been normalized. The similarity calculation unit 121 calculates the similarity using the frequency distribution based on the data after the frequency analysis. FIG. 4 shows data Q1 to Q3 obtained by performing frequency analysis on data of different time frames. The vertical axis of the graph in FIG. 4 indicates the magnitude of the frequency components, and the horizontal axis indicates the frequency. For example, the value on the vertical axis of data Q1 indicates the average value (average in the time axis direction) of the magnitudes of the respective frequency components in time frame W1. When comparing data Q1 to Q3, even though the magnitudes of the frequency components are different, the waveform shapes are similar. For example, in any of data Q1 to Q3, in the range of frequency f1, as the frequency increases, the magnitude of the frequency component increases, and in the range of f2, even if the frequency changes, the magnitude of the frequency component is generally constant. When the trends of the frequency distributions are similar in this way, the similarity between data Q1 to Q3 is high, and when the trends of the frequency distributions are different, the similarity is low. Specifically, the similarity calculation unit 121 calculates, for example, the cosine similarity between vectors V1 and V3 composed of the respective frequency components of data Q1 and data Q3 in the frequency range f3. The similarity calculation unit 121 also calculates the cosine similarity between vectors (for example, V1a and V3a) composed of the respective frequency components in the corresponding range for other frequency ranges. For example, the similarity calculation unit 121 sets the average of the cosine similarities calculated for a predetermined range for each of variables A to C as the similarity between data Q1 and data Q3. More specifically, the similarity calculation method using the fast Fourier transform is as follows. (Step 1) Perform fast Fourier transform on two pieces of data to calculate the real and imaginary numbers of the components of each frequency. (Step 2) Calculate the cosine similarity using the real and imaginary coefficients calculated above as features as the similarity.

[0020] The specified range is, for example, when the data contains a lot of noise (mainly high-frequency components), the range of other frequencies excluding the frequency range (high frequency) where noise is abundantly contained. Also, when the frequency component value is 0 or a low value for a certain frequency range, the cosine similarity may be calculated after excluding that range. Further, for example, depending on the operating mode of the target device or plant (during startup, rated operation, partial load operation, stopped, etc.), the similarity may be calculated using only the frequency range corresponding to each operating mode. For example, when in a transient operating state (during startup, stopped, or changing the operating mode), the similarity may be calculated using the data in the low-frequency range. The similarity calculation unit 121 calculates the similarity based on the frequency distributions of the time frame W1 and each of the time frames W2 to W9 in such a method, calculates the similarity between the time frame W2 and each of the time frames W3 to W9, and calculates the similarity for all patterns of combinations of time frames in the same way. Note that in the similarity based on the frequency component distribution, the temporal front-back relationship is originally excluded, and the feature amount (frequency component value of each frequency) and the similarity are calculated.

[0021] Based on the idea that the characteristics of certain data can be replaced by other data with high similarity, the data extraction unit 12 uses the data reduction unit 122 to reduce the data with overlapping time-series characteristics and extracts some data with the characteristics necessary for building the learning model, based on the similarity calculated by the similarity calculation unit 121. Next, the data reduction process will be described with reference to FIGS. 5A to 5D.

[0022] [Data Reduction Process] Table 500 in FIG. 5A shows the result of organizing the similarity between each data of A to I cut out in each time frame by the above similarity calculation method of (1) or (2). For example, the similarity between data A and data B is 0.9, the similarity between data A and data C is 0.6, and the similarity between data A and data I is 0.2. Here, the data reduction unit 122 classifies each data into similar data and non-similar data based on a predetermined threshold value (for example, 0.9). The group of similar data after classification is shown in Table 501 in FIG. 5B. In addition, FIG. 5C shows the distribution based on the feature amount indicating the similarity between each data by the distance between the data. Note that data A to I belong to class 1, and data A' belongs to class 2. For example, class 1 is a set of data of each variable such as pressure and flow rate when the state of the plant is normal, and class 2 is a set of data when the state of the plant is abnormal. Alternatively, class 1 is a set of data of each variable when the plant is operating at rated load, and class 2 is a set of data at startup. The data reduction unit 122 deletes learning data within each class. Here, the data deletion within class 1 will be described.

[0023] The data reduction unit 122 selects data having a large number of highly similar data and deletes the selected data. Referring to Table 501 in Fig. 5B, for example, the number of data in class 1 similar to data A is 1 (B), the number of data in class 1 similar to data C is 3 (G, H, I), and the number of data in class 1 similar to data H is 5 (C, E, F, G, I), and so on. Also, referring to Fig. 5C, it can be seen that data C, E, F, G, I with a similarity of 0.9 or more are distributed around data H. In such a case, the data reduction unit 122 deletes data H. The relationship between the data after deletion is shown in Table 502 in Fig. 5B. Referring to Table 502, the number of data similar to data B is 2 (A, C), the number of data similar to data C is 2 (B, I), and for other data, the number of similar data is 1 or less. For the data group in Table 502, the data reduction unit 122 repeatedly performs the process of deleting data having a large number of highly similar data. For example, the data reduction unit 122 deletes either data B or data C (either is fine). Here, it is assumed that data B is deleted. The relationship between the data after deleting data B is shown in Fig. 5D. In Fig. 5D, the number of data similar to each data is 1 or less. By repeatedly performing such a process, data with overlapping time-series features is deleted. The data reduction unit 122 ends the data reduction when it can sufficiently reduce the number of learning data while preserving the features of the original learning data group. For example, the data reduction unit 122 ends the deletion of data when the total number of data becomes equal to or less than a predetermined number, or when the number of data similar to each data becomes equal to or less than a predetermined number. The data extraction unit 12 extracts the data remaining after reducing the data with overlapping time-series features as learning data for learning the rough structure of the learning model.

[0024] (Operation) [Data Extraction Process] Next, referring to Fig. 6, the flow of the data extraction process of the present embodiment will be described. Fig. 6 is a flowchart showing an example of the data extraction process of the first embodiment. First, the data acquisition unit 11 acquires a plurality of pieces of learning data cut out for each time window from the time-series data of a plurality of parameters (for example, pressure, flow rate, ···, etc.) (step S10). The data acquisition unit 11 may acquire the time-series data, cut out the data for each time window, and acquire the learning data. The data acquisition unit 11 records the plurality of pieces of learning data cut out for each time frame in the storage unit 16.

[0025] Next, the data extraction unit 12 classifies the learning data into classes (step S20). For example, when it is desired to divide the learning data into data (class 1) when the plant is in a normal operating state and data (class 2) when it is abnormal, the user inputs the time period of the normal operating state and the time period of the abnormal operating state to the learning model construction device 10. The input unit 14 outputs the information of the input time period to the data extraction unit 12. The data extraction unit 12 classifies the learning data acquired in step S10 into either class 1 or class 2 according to the time period set by the user.

[0026] Next, the data extraction unit 12 performs a process (steps S30 to S60) of excluding data with overlapping time-series features from the plurality of pieces of learning data and extracting only a small number of pieces of learning data necessary for accurately constructing the rough structure of the learning model. The data extraction unit 12 executes this process for each class.

[0027] First, the similarity calculation unit 121 calculates (1) the similarity based on the normalized cross-correlation or (2) the similarity based on the frequency distribution (step S30). The similarity calculation unit 121 calculates the similarity by the method of (1) or (2) for all combination patterns of selecting and combining two pieces of learning data belonging to the same class classified in step S20, and records the similarity between each piece of learning data in the storage unit 16. As a result, data such as the table 500 in FIG. 5A is obtained. The similarity calculation unit 121 may calculate the similarity by both methods of (1) and (2) and use their weighted average as the similarity.

[0028] Next, the data reduction unit 122 aggregates combinations of highly similar learning data (step S40). For each learning data, the data reduction unit 122 associates other learning data that is equal to or higher than a predetermined threshold (for example, 0.9 or higher), and separates data with high similarity from each other. The data reduction unit 122 records this separation result in the storage unit 16. By this process, data as exemplified in Table 501 of FIG. 5B is obtained.

[0029] Next, the data reduction unit 122 deletes learning data having a large number of highly similar data (step S50). When learning data as exemplified in FIG. 5C is obtained, the data reduction unit 122 deletes data H. By deleting data in this way, it is possible to reduce the number of learning data while retaining the features of the learning data group (while suppressing the loss of features). In the case of the data distribution exemplified in FIG. 5C, by leaving data C, I, G, F, etc. and deleting data H, it is possible to reduce the data amount while leaving learning data scattered near the boundary of class 1. As another method of reducing learning data, learning data having a large number of highly similar data may be selected, the selected data may be left, and other similar learning data may be deleted. For example, in the case of Table 501 of FIG. 5B, regarding data H, data H is left and data C, E, F, G, I are deleted. As a result, in the case of the example of FIG. 5C, more data can be reduced.

[0030] Next, the data reduction unit 122 determines whether or not to complete the deletion of the learning data (step S60). For example, when the number of data similar to each data is equal to or less than a predetermined number, the data reduction unit 122 determines to complete the deletion of the learning data. Alternatively, when the number of remaining learning data is equal to or less than a predetermined number, or when the cumulative number of deleted data reaches a predetermined number, the data reduction unit 122 may determine to complete the deletion of the learning data. If the deletion of the learning data is not completed (step S60; No), the process from step S40 is repeated. If the deletion of the learning data is completed (step S60; Yes), the data extraction unit 12 records the extracted learning data in the storage unit 16 (step S70). The data extraction unit 12 records the remaining learning data that has not been deleted as the extracted learning data in the storage unit 16.

[0031] [Learning model construction process] Next, the process of constructing a learning model will be described. FIG. 7 is a flowchart showing an example of the learning model construction process of the first embodiment. First, the data acquisition unit 11 acquires learning data cut out for each time frame from the time series data of a plurality of parameters (step S1). Next, the data extraction unit 12 extracts the learning data (step S2). The data extraction unit 12 extracts the learning data by the process described in FIG. 6.

[0032] Next, the learning unit 13 learns the data extracted by the data extraction unit 12 (step S3). The learning unit 13 reads out the extracted learning data from the storage unit 16 and constructs a learning model by deep learning such as CNN, for example. For example, the learning unit 13 learns teacher data in which information indicating the plant state (normal, abnormal, occurrence of event XX, etc.) in the time zone of the time window is labeled for the learning data of a plurality of parameters of a common time window, and constructs a discrimination model for discriminating events occurring in the plant. As a result, approximate values of the hyperparameters of the discrimination model are determined. In this step, since the learning data is narrowed down, learning and evaluation can be performed in a shorter time compared to the case where learning is performed using all the learning data. Also, since the data extracted by the process of FIG. 6 retains the characteristics of the original learning data group, a learning model (discrimination model) can be constructed without degrading the accuracy.

[0033] Next, the learning unit 13 learns all the learning data (step S4). The learning unit 13 advances learning using a large number of learning data for the learning model constructed in step S3 and adjusts the determined values of the hyperparameters. Since the values of the hyperparameters have already been determined in step S3, learning and evaluation can be performed in a relatively short time even when learning is performed using all the learning data. Also, by performing learning using a large number of data, the accuracy of the learning model constructed in step S3 can be improved. According to the process of FIG. 7, by performing the process of step S3 before constructing the learning data using all the learning data, the learning model can be constructed in a shorter time compared to the case where the learning model is constructed using all the learning data from the beginning.

[0034] (Effect) As described above, according to the present embodiment, it is possible to reduce the number of learning data while retaining the characteristics of the original learning data group. By reducing the learning data while maintaining the accuracy, it is possible to shorten the learning time required for building the learning model. Assuming that the structure of the deep learning model depends not on the amount of learning data but on the variation in the feature amounts of the learning data, after specifying the model structure that can obtain the desired accuracy using the learning data extraction method of the present embodiment, by using a large amount of learning data and fine-tuning each parameter to improve the accuracy of the model, it is possible to efficiently build the learning model.

[0035] The learning data extraction method of the present embodiment is effective when it is not possible to attach a learnable model or a high-precision model structure and it is necessary to evaluate while changing the model structure several times. Also, the learning data extraction method of the present embodiment is effective for learning methods (such as deep learning) that ignore the temporal context of the features of the learning data (exclude temporal information) when learning time-series data. Further, for example, it can be applied to a system that deals with a multi-valued classification problem of time-series data, such as a discrimination model for discriminating abnormal events in a plant. In the embodiment, an example of calculating the similarity using normalized cross-correlation and frequency distribution is given to determine whether the data cut out from the time window has the same features in time series, but the similarity can also be calculated by other methods. Also, in the above embodiment, supervised deep learning is used as an example for explanation, but it can also be applied to unsupervised deep learning.

[0036] <Second Embodiment> Hereinafter, the learning model construction apparatus 10A according to the second embodiment of the present invention will be described with reference to FIGS. 8 to 9. In the first embodiment, the similarity was calculated for all combinations of learning data. This method may result in an enormous computational load and calculation time when the number of learning data becomes extremely large. Therefore, in the second embodiment, the calculation of the similarity between data with low similarity is omitted to reduce the amount of calculation.

[0037] (Configuration) FIG. 8 is a block diagram showing an example of the learning model construction device according to the second embodiment. Among the configurations of the learning model construction device 10A according to the second embodiment of the present invention, those that are the same as the functional units constituting the learning model construction device 10 according to the first embodiment of the present invention are denoted by the same reference numerals, and their descriptions are omitted. The learning model construction device 10A according to the second embodiment includes a data extraction unit 12A instead of the data extraction unit 12 in the first embodiment. The data extraction unit 12A includes a classification unit 123 in addition to the similarity calculation unit 121 and the data reduction unit 122 similar to those in the first embodiment.

[0038] The classification unit 123 classifies the learning data of the same class into groups (clusters) in which similar ones are gathered. For example, the classification unit 123 classifies the learning data whose (1) normalized cross-correlation or (2) trend of frequency component distribution is similar, as described in the first embodiment, into the same cluster by an arbitrary clustering method. The similarity calculation unit 121 calculates the similarity for combinations of learning data belonging to the same cluster.

[0039] (Operation) Next, the data extraction process of the second embodiment will be described. FIG. 9 is a flowchart showing an example of the data extraction process of the second embodiment. For the same processing as that of the first embodiment described with reference to FIG. 6, the same reference numerals are used and the description thereof is omitted. First, the data acquisition unit 11 acquires a plurality of learning data (step S10). Next, the data extraction unit 12 classifies the learning data into classes (step S20). Next, the classification unit 123 performs clustering on the learning data within the same class (step S25). The classification unit 123 classifies similar learning data into the same cluster for the learning data belonging to class 1 based on an arbitrary clustering method (for example, the kmeans method) and a predetermined evaluation measure (for example, the similarity of feature amounts not considering the temporal context between learning data). For example, the classification unit 123 calculates the similarity based on (1) the normalized cross-correlation or (2) the frequency component distribution, and classifies those with high similarity into one of a plurality of clusters. The classification unit 123 performs clustering for each class. However, if the number of learning data in a certain class is less than the threshold value, the processing of step S31 may not be performed for that class. The classification unit 123 records the identification information of the cluster to which the learning data belongs in the storage unit 16 for all the learning data for which clustering has been performed.

[0040] Next, the similarity calculation unit 121 calculates the similarity for all combination patterns in which two are selected and combined from the learning data classified into the same cluster by the methods of (1) and (2) (step S31), and records the similarity between each learning data in the storage unit 16. If the similarity has been calculated in the same manner in step S25, the calculation result can be used. By limiting the calculation of the similarity within the range of the same cluster of the same class, the calculation cost can be reduced as compared with step S30 of the first embodiment.

[0041] Next, the data reduction unit 122 aggregates combinations of highly similar learning data (step S40). In step S25, since the learning data that are similar are classified based on the feature amounts that do not consider the temporal context, even if the target range of the similarity calculation is limited in step S31, there is no influence on the aggregation result in step S40 (for example, in the table 501 of FIG. 5B, data with high similarity cannot be separated, etc.), and it is considered that the extraction accuracy of the learning data does not decrease. Next, the data reduction unit 122 deletes the learning data having a large number of highly similar data (step S50). The data reduction unit 122 repeats the data reduction until a predetermined completion condition is satisfied (step S60). The data extraction unit 12 records the learning data that remains without being deleted in the storage unit 16 as the extracted learning data (step S70). The process of constructing a learning model using the extracted learning data is the same as that in the first embodiment (FIG. 7).

[0042] According to the second embodiment, when the combinations between learning data become enormous, clustering is performed, and the time required for similarity calculation is reduced by omitting the similarity calculation with the learning data belonging to other clusters considered to have low similarity. Thereby, in addition to the effect of the first embodiment, the time for the data extraction process can be shortened. In the clustering by the kmeans method, in order to estimate the cluster to which each learning data belongs, it is necessary to construct a kmeans model. When constructing a kmeans model, all the learning data may be used, but if it takes time to construct the model, the kmeans model may be constructed using representative learning data.

[0043] FIG. 10 is a diagram showing an example of the hardware configuration of the learning model construction device according to each embodiment. The computer 900 includes a CPU 901, a main storage device 902, an auxiliary storage device 903, an input / output interface 904, and a communication interface 905. The above-described learning model construction apparatuses 10 and 10A are implemented in a computer 900. And each of the above-described functions is stored in an auxiliary storage device 903 in the form of a program. The CPU 901 reads the program from the auxiliary storage device 903 and expands it in the main storage device 902, and executes the above processing according to the program. Further, the CPU 901 secures a storage area in the main storage device 902 according to the program. Also, the CPU 901 secures a storage area in the auxiliary storage device 903 for storing data being processed according to the program.

[0044] Note that a program for realizing all or part of the functions of the learning model construction apparatuses 10 and 10A may be recorded on a computer-readable recording medium, and the program recorded on this recording medium may be read into a computer system and executed to perform processing by each functional unit. Here, the “computer system” is assumed to include hardware such as an OS and peripheral devices. Also, the “computer system” includes a homepage providing environment (or display environment) if it uses a WWW system. Also, the “computer-readable recording medium” refers to a portable medium such as a CD, DVD, USB, or a storage device such as a hard disk built into a computer system. Also, when this program is distributed to the computer 900 via a communication line, the computer 900 that has received the distribution may expand the program in the main storage device 902 and execute the above processing. Also, the above program may be for realizing a part of the above-described functions, and may further be realized in combination with a program already recorded in the computer system for the above-described functions. Note that the learning model construction apparatuses 10 and 10A may be configured by a plurality of computers 900.

[0045] As described above, some embodiments according to the present disclosure have been explained. However, all of these embodiments are presented as examples and are not intended to limit the scope of the invention. These embodiments can be implemented in various other forms, and various omissions, replacements, and changes can be made without departing from the gist of the invention. These embodiments and their modifications are included in the scope and gist of the invention, as well as in the invention described in the claims and the equivalent scope thereof.

[0046] <Appendix> The data extraction device, learning model construction device, data extraction method, learning model construction method, and program described in each embodiment are understood as follows, for example.

[0047] (1) The data extraction device (learning model construction device 10) according to the first aspect is a data extraction device that extracts some learning data from a plurality of learning data cut out for each time window from time series data, and includes a data acquisition unit 11 that acquires the plurality of learning data, and a similarity calculation unit 121 that calculates the similarity between the learning data based on the characteristics of the learning data, excluding the temporal front-back difference in the time window, and a data reduction unit 122 that separates the learning data with high similarity and deletes a part of them. Thereby, in the learning method (such as deep learning), while ignoring the temporal front-back relationship of the characteristics of the learning data (excluding time series information), the number of learning data can be effectively reduced while maintaining the characteristics of the original learning data group.

[0048] (2) The data extraction device (learning model construction device 10) according to the second aspect is the data extraction device of (1), and the data reduction unit 122 associates other learning data with a similarity equal to or higher than a threshold value for each of the learning data, and deletes the learning data with the most associated other learning data. Thereby, the learning data can be reduced without impairing the characteristics of the learning data group.

[0049] (3) The data extraction device (learning model construction device 10) according to the third aspect is the data extraction device of (1) to (2), and the similarity calculation unit 121, when shifting the two pieces of learning data in the time axis direction, calculates the similarity based on the magnitude per unit time of the area where the regions related to the two pieces of learning data overlap with respect to the region sandwiched between the waveform data indicated by each piece of learning data and the time axis. Thereby, the similarity between learning data can be calculated based on the feature amount excluding the temporal front-back difference.

[0050] (4) The data extraction device (learning model construction device 10) according to the fourth aspect is the data extraction device of (1) to (3), and the similarity calculation unit 121 calculates the similarity between the two shapes based on the shape of each frequency distribution obtained by performing frequency analysis on the two pieces of learning data. Thereby, the similarity between learning data can be calculated based on the feature amount independent of time (the feature amount excluding the temporal front-back difference).

[0051] (5) The data extraction device (learning model construction device 10) according to the fifth aspect is the data extraction device of (1) to (4), and further includes a classification unit 123 that classifies a plurality of the learning data into the same group (cluster) based on a predetermined evaluation scale (for example, the similarity of feature amounts not considering the temporal front-back relationship between learning data), and the similarity calculation unit 121 calculates the similarity between the learning data belonging to the same group (cluster). Thereby, it is possible to suppress the calculation cost of similarity calculation from becoming enormous.

[0052] (6) The data extraction device (learning model construction device 10) according to the sixth aspect is the data extraction device of (1) to (5), and the learning data is learning data used for deep learning. Thereby, the amount of learning data for deep learning can be reduced. The reduced learning data can be used to determine the general structure of the deep learning model.

[0053] (7) The learning model construction device 10 according to the seventh aspect includes the data extraction device of (1) to (6) and a learning unit that learns the learning data to construct a learning model. After constructing the learning model with the learning data extracted by the data extraction device, the learning unit adjusts the parameters of the learning model using the learning data before extraction. Thereby, the time required for constructing the learning model can be reduced. In addition, by adjusting the parameters of the learning model using all the data, the accuracy of the learning model can be ensured.

[0054] (8) The data extraction method according to the eighth aspect is a data extraction method for extracting some of the plurality of learning data cut out from the time-series data for each time window. The method includes steps of: acquiring the plurality of learning data; calculating the similarity between the learning data based on the features of the learning data by excluding the temporal differences before and after in the time window; and separating the learning data with high similarity from each other and deleting a part of them.

[0055] (9) The learning model construction method according to the ninth aspect constructs the learning model by learning the learning data extracted by the data extraction method of (8), and then adjusts the parameters of the learning model using the learning data before extraction.

[0056] (10) The program according to the tenth aspect causes a computer 900 to execute a data extraction process for extracting some of the plurality of learning data cut out from the time-series data for each time window. The data extraction process includes steps of: acquiring the plurality of learning data; calculating the similarity between the learning data based on the features of the learning data by excluding the temporal differences before and after in the time window; and separating the learning data with high similarity from each other and deleting a part of them.

[0057] (11) The program according to the eleventh aspect causes the computer 900 to execute a process of adjusting the parameters of the learning model using the learning data before extraction, after building the learning model by learning the learning data extracted by the data extraction process of (8).

Explanation of Signs

[0058] 10, 10A... Learning model construction device 11... Data acquisition unit 12, 12A... Data extraction unit 121... Similarity calculation unit 122... Data reduction unit 123... Classification unit 13... Learning unit 14... Input unit 15... Output unit 16... Storage unit 900... Computer 901... CPU 902... Main memory device 903... Auxiliary storage device 904... Input / output interface 905... Communication interface

Claims

1. A data extraction device that extracts some learning data from a plurality of learning data cut out for each time window from time-series data, comprising: a data acquisition unit that acquires a plurality of the learning data; a similarity calculation unit that calculates the similarity between the learning data based on the features of the learning data excluding the temporal differences before and after in the time window; a data reduction unit that separates the learning data with high similarity and deletes a part of them; wherein the similarity calculation unit calculates the similarity based on the magnitude per unit time of the overlapping area of the regions related to the two learning data with respect to the region sandwiched between the waveform data indicated by each learning data when the two learning data are shifted in the time axis direction; a data extraction device.

2. A data extraction device that extracts some learning data from a plurality of learning data cut out for each time window from time-series data, comprising: a data acquisition unit that acquires a plurality of the learning data; a similarity calculation unit that calculates the similarity between the learning data based on the features of the learning data excluding the temporal differences before and after in the time window; a data reduction unit that separates the learning data with high similarity and deletes a part of them; wherein the similarity calculation unit calculates the similarity between the two shapes based on the shape of each frequency distribution obtained by frequency analyzing the two learning data; a data extraction device.

3. A data extraction device that extracts some learning data from a plurality of learning data cut out for each time window from time-series data, comprising: a data acquisition unit that acquires a plurality of the learning data; a classification unit that classifies the learning data similar to each other into the same group based on a predetermined evaluation criterion; a similarity calculation unit that calculates the similarity between the learning data based on the features of the learning data excluding the temporal differences before and after in the time window; a data reduction unit that separates the learning data with high similarity and deletes a part of them; wherein the similarity calculation unit calculates the similarity between the learning data belonging to the same group; a data extraction device.

4. The data reduction unit associates other learning data with a threshold value or more of similarity for each of the learning data, and deletes the learning data with the most associated other learning data. The data extraction device according to any one of claims 1 to 3.

5. The learning data is learning data used for deep learning. The data extraction device according to any one of claims 1 to 4.

6. A data extraction device according to any one of claims 1 to 5, and a learning unit that learns the learning data to construct a learning model. The learning unit learns the learning data extracted by the data extraction device to construct the learning model, and then adjusts the parameters of the learning model using the learning data before extraction. Learning model construction device.

7. A data extraction method for extracting some learning data from a plurality of learning data cut out for each time window from time series data, comprising the steps of: obtaining a plurality of the learning data; calculating the similarity between the learning data based on the features of the learning data excluding the temporal differences before and after in the time window; separating the learning data with high similarity from each other and deleting a part of them; and having In the step of calculating the similarity, when two pieces of the learning data are shifted in the time axis direction, the similarity is calculated based on the magnitude per unit time of the overlapping area of the regions sandwiched by the waveform data indicated by each learning data and the time axis for the regions related to the two pieces of the learning data. Data extraction method.

8. A data extraction method for extracting some learning data from a plurality of learning data cut out for each time window from time series data, comprising the steps of: obtaining a plurality of the learning data; calculating the similarity between the learning data based on the features of the learning data excluding the temporal differences before and after in the time window; separating the learning data with high similarity from each other and deleting a part of them; and having In the step of calculating the similarity, the similarity between the two shapes is calculated based on the shape of each frequency distribution obtained by frequency analyzing the two pieces of the learning data. Data extraction method.

9. A data extraction method for extracting some learning data from a plurality of learning data cut out for each time window from time series data, comprising the steps of: obtaining a plurality of the learning data; classifying the plurality of the learning data into the same group based on a predetermined evaluation scale for those that are similar to each other. Calculating the similarity between the learning data based on the features of the learning data excluding the temporal differences before and after in the time window; Separating the learning data with high similarity and deleting a part of them; comprising: In the step of calculating the similarity, the similarity between the learning data belonging to the same group is calculated. Data extraction method.

10. Training a learning model by training the learning data extracted by the data extraction method according to any one of Claims 7 to 9, and then adjusting the parameters of the learning model using the learning data before extraction. Learning model construction method.

11. On a computer, A data extraction process for extracting some learning data from a plurality of learning data cut out for each time window from time series data, Obtaining a plurality of the learning data; Calculating the similarity between the learning data based on the features of the learning data excluding the temporal differences before and after in the time window; Separating the learning data with high similarity and deleting a part of them; comprising: In the step of calculating the similarity, when two pieces of the learning data are shifted in the time axis direction, the similarity is calculated based on the magnitude per unit time of the overlapping area of the regions sandwiched by the waveform data and the time axis indicated by each learning data. A program for executing the data extraction process.

12. On a computer, A data extraction process for extracting some learning data from a plurality of learning data cut out for each time window from time series data, Obtaining a plurality of the learning data; Calculating the similarity between the learning data based on the features of the learning data excluding the temporal differences before and after in the time window; Separating the learning data with high similarity and deleting a part of them; comprising: In the step of calculating the similarity, a program for executing a data extraction process for calculating the similarity between two shapes based on the shape of each frequency distribution obtained by frequency analyzing two pieces of the learning data is executed.

13. On a computer, A data extraction process for extracting some learning data from a plurality of learning data cut out for each time window from time series data, The step of acquiring a plurality of the learning data; The step of classifying similar ones among the plurality of the learning data into the same group based on a predetermined evaluation measure; The step of calculating the similarity between the learning data based on the features of the learning data excluding the temporal front-back difference in the time window; The step of separating the learning data with high similarity from each other and deleting a part thereof; comprising; In the step of calculating the similarity, a program for executing a data extraction process of calculating the similarity between the learning data belonging to the same group. [

14. ] On a computer, A program for causing the computer to learn the learning data extracted by the program according to any one of Claims 11 to 13 to construct a learning model, and then execute a process of adjusting parameters of the learning model using the learning data before extraction.

Citation Information

Patent Citations

  • Target detection performance optimization method

    CN106934346A

  • Neural network learning method

    JP1994004506A

  • Method and device for learning example

    JP1994348678A

  • Representative information selection method, representative information selection system and program

    JP2008071136A

  • Recurrent type neural network learning method, computer program for the same, and voice recognition device

    JP2016212273A