Data processing system and data processing method

The data processing system improves predictive maintenance by identifying and managing impact data within training sets, enhancing detection accuracy through unsupervised learning and user-controlled data selection.

JP7767181B2Active Publication Date: 2025-11-11HIATACHI POWER SOLUTIONS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022027097
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-02-24
Publication Date
2025-11-11
Estimated Expiration
2042-02-24

AI Technical Summary

Technical Problem

Conventional systems face challenges in selecting suitable training data for predictive maintenance models due to the arbitrary nature of data selection criteria and the inclusion of unrecorded failures or abnormal data, which affects detection accuracy.

Method used

A data processing system that utilizes unsupervised learning to identify and prioritize impact data within training data, allowing users to decide whether to include or exclude this data for improved model construction.

Benefits of technology

Enhances the detection accuracy of predictive maintenance models by assisting in the selection of training data, ensuring that only relevant data is used for machine learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007767181000001
    Figure 0007767181000001
  • Figure 0007767181000002
    Figure 0007767181000002
  • Figure 0007767181000003
    Figure 0007767181000003
Patent Text Reader

Abstract

To improve the detection accuracy of a model by assisting in screening learning data upon creating a model to detect a sign of abnormality of an object by performing machine learning based on data measured from the object.MEANS FOR SOLVING THE PROBLEM: A system for processing data includes a controller that receives time-series data from a measurement sensor of an industrial machine and executes a program, thereby executing data processing for detecting a sign of abnormality of an object based on the time-series data. The controller gathers learning data for machine learning from the received time-series data, the learning data including a plurality of measurement data, performs machine learning based on the learning data and constructs a model for detecting a sign of abnormality of the industrial machine, calculates an abnormality degree of verification data by applying the model to verification data, identifies measurement data exerting an influence upon an abnormality degree out of the measurement data included in the learning data as data to be noticed, and presents the data to be noticed.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a data processing technique for detecting signs of abnormality occurring in an object, and more particularly to an invention for data processing for detecting signs of failure in industrial machinery. [Background technology]

[0002] An evaluation model is constructed based on time-series data measured from the object of analysis, and the evaluation model is used to evaluate the condition of the object of analysis. For example, Condition Based Maintenance (CBM) is known as a maintenance method for maintenance targets such as facilities, machinery, equipment, and devices.

[0003] This involves monitoring the condition of equipment, even if it is currently operating stably, and maintaining it before it breaks down depending on the state of deterioration. For example, the operating conditions of equipment, such as vibration and heat generation in wind power generation facilities and gas turbines, are monitored, and failure predictions are made based on the results.

[0004] A predictive maintenance system creates a learning model using machine learning based on time-series data obtained from equipment and facilities, and then applies time-series data from sensors to this model to predict equipment failures. The system takes in normal data from the sensor data as learning data, builds a detection model based on the learning data, and verifies the model's detection performance by applying the equipment's past failure history as verification data.

[0005] In this type of system, it is desirable to improve the suitability of the training data in order to improve the performance of the evaluation model. For example, JP 2020-102001 A discloses a training data confirmation support device that uses machine learning to detect abnormalities in industrial machinery, with the aim of making it easier to confirm whether the data was measured using the same operation when acquiring measurement data from industrial machinery. The device makes it easier to check for the inclusion of inappropriate data when acquiring training data consisting only of normal data in advance, and determines whether inappropriate data is mixed in by matching it with past data when generating a training model.

[0006] Furthermore, Patent Publication No. 2021-33554 discloses a method for refining training data that improves the predictive accuracy of a model by identifying harmful training data for each of multiple training data sets based on a score that represents the strength of the impact that the harmful training data has on the predictive accuracy of the model for one sample data set, and deleting the harmful training data from the training data. [Prior art documents] [Patent documents]

[0007] [Patent Document 1] Japanese Patent Publication No. 2020-102001 [Patent Document 2] Patent Publication No. 2021-33554 Summary of the Invention [Problem to be solved by the invention]

[0008] Conventional systems select normal data from sensor data from periods when the equipment is not malfunctioning, and use the selected data as training data to generate a predictive maintenance model for the equipment. However, determining whether data is suitable for use as training data is not easy in the first place. For example, even data from periods when the equipment is not malfunctioning may contain unrecorded failures, and the equipment may be operating abnormally without clearly indicating a failure. Conventionally, this type of data has been excluded from the training data based on predetermined rules, but past rules are often no longer valid for new equipment or new operating environments.

[0009] On the other hand, if the selection of training data is left to humans, when the period of time-series data is long, the number of data selection patterns becomes enormous, imposing a heavy burden, and the selection criteria become arbitrary and lack versatility. Therefore, an object of the present invention is to improve the detection accuracy of a model by supporting the selection of training data when performing machine learning based on data measured from a target to create a detection model for signs of abnormality in the target. [Means for solving the problem]

[0010] To achieve the above object, the present invention provides a data processing system including a controller that receives time-series data from a measurement sensor of an object and executes a program to perform data processing based on the time-series data to detect signs of anomaly in the object, wherein the controller collects learning data for machine learning from the received time-series data, the learning data includes a plurality of measurement data, performs machine learning based on the learning data to construct a model for detecting signs of anomaly in the object, applies the model to validation data to calculate an anomaly level for the validation data, identifies measurement data included in the learning data that has an effect on the anomaly level as data that should be noted, and presents the data that should be noted. The present invention also provides a data processing method according to this feature. [Effects of the Invention]

[0011] According to the present invention, when machine learning is performed based on data measured from a target to create a detection model for predicting abnormalities in the target, the detection accuracy of the model can be improved by assisting in the selection of training data. [Brief explanation of the drawings]

[0012] [Figure 1] 1 is a block diagram of an overall system that includes an embodiment of a data processing system and detects a sign of an abnormality in a target. [Figure 2] 10 is an example of a recording table of sensor data. [Figure 3] 10 is an example of a failure history table. [Figure 4] FIG. 4 is a waveform diagram showing an example of time-series data output from a measurement sensor of the device. [Figure 5] 10 is a flowchart illustrating an example of a data processing program. [Figure 6A] 10 is a waveform diagram for explaining division of time-series data output from a measurement sensor of the device at predetermined time intervals. FIG. [Figure 6B] FIG. 10 is a diagram for explaining how time-series data output from a measurement sensor of the device is divided by clustering. [Figure 6C] FIG. 10 is a diagram showing the relationship between division of time series data by clustering and a waveform diagram of the time series data. [Figure 7] This is a block that explains the process of identifying impact data (data that should be noted) from data for machine learning. [Figure 8] 10 is a graph showing the change over time of the target abnormality degree and the reference abnormality degree. [Figure 9A] 10 is another graph showing the change over time of the target abnormality degree and the reference abnormality degree. [Figure 9B] 10 is yet another graph showing the change over time of the target abnormality degree and the reference abnormality degree. [Figure 9C] 10 is yet another graph showing the change over time of the target abnormality degree and the reference abnormality degree. [Figure 10] FIG. 10 is a block diagram illustrating how the data refinement module determines the priority of a data block. [Figure 11] 10 is an example of a display list as one of the information display modes of impact data. [Figure 12] 10 is an example of a screen showing impact data related information output to a user device. DETAILED DESCRIPTION OF THE INVENTION

[0013] Next, an embodiment of a data processing system according to the present invention will be described. Fig. 1 is a hardware block diagram of an overall system 100 that detects signs of abnormality in an object, including this data processing system 101. The object may be an industrial machine that requires detection of signs of failure, such as a wind power generation device or a gas turbine engine.

[0014] The overall system 100 comprises a device group 110 having a plurality of devices 111 each as a target, a sensor group 120 consisting of sensors 121 (sensor 1 to sensor N) of each of the plurality of devices, storage 130 that stores sensor data 131, which is measurement data from the sensors, and device failure history data 132, and a data processing system 101.

[0015] The data processing system 101 includes a data refining server (data refining device) 150 and a failure prediction server (failure sign detection device) 160 that detects signs of failure. The device group 110, sensor group 120, storage 130, data refining server 150, and failure prediction server 160 are connected to one another via a network 165.

[0016] When the storage 130 receives sensor data from the sensor group, it records measurement data 203, 204 for each device ID 201 of the device group and for each sensor of the sensor group in the sensor data (storage area) 131, as shown in Fig. 2 (sensor data recording table). 202 is the measurement time. The sensor outputs time series data measured from the device, such as time series data of control amount related to device control, time series data of state amount related to device control, and time series data of the environment in which the device is operating.

[0017] Then, when the storage 130 receives data related to the failure from the device, as shown in Fig. 3 (failure history table), it records, for each failure ID 301, the device ID 302 on which the failure occurred, the failure date 303 on which the failure occurred, the symptoms of the failure 304, and the countermeasures taken against the failure 305 in a table of failure history (storage area) 132. As shown in Fig. 4, the data processing system 101 sets normal data that does not contain abnormal data such as output data at the time of device failure 406 from the output data 400 of the sensor of the device 111 as learning data 402, in other words, data that is only normal data or mainly normal data, and performs machine learning based on the learning data 402 to construct a detection model for signs of failure of the device.

[0018] Because the distinction between normal data and abnormal data is relative depending on the state of the device and the operating environment of the device, the data processing system 101 preferably performs unsupervised learning based on the sensor data to create a learning model. The data processing system 101 applies the failure sign detection model to validation data 404 including sensor data from before the device failure 406, calculates the degree of anomaly of the validation data 404, and obtains an evaluation index for the detection performance of the failure sign detection model.

[0019] The data refining server 150 includes a data refining module 151 for optimizing learning data collected from sensor data, and a model construction module 152 for performing machine learning based on the optimized learning data to construct a failure sign detection model for the device. The data refining server 150 realizes the data refining module 151 and the model construction module 152 by having a controller execute a program stored in memory. Note that the module may also be referred to as a means, a part, or a unit. The failure prediction server 160 also realizes a module 161 for applying a failure sign detection model to sensor data by having a controller execute a program stored in memory.

[0020] The data refinement module 151 provides support processing for optimizing the training data 402 to the administrator, user, etc. of the data processing system 101. Within the system of sensor data collected as training data 402, the data refinement module 151 relatively identifies data that has a predetermined or greater impact on the accuracy of failure sign detection, and presents this to the user as impact data, i.e., data that should be noticed by the user, along with its degree of impact.

[0021] Upon receiving this presentation, the user can decide whether to perform machine learning while keeping the impact data included in the learning data, or to perform machine learning after excluding the impact data from the learning data, and the model construction module 152 will perform machine learning using either method based on instructions from the user.

[0022] The operation of data processing system 101 will be explained based on the flowchart in Fig. 5. The flowchart is executed by a processor or controller that executes a program to realize data refinement module 151, etc. To clarify the relationship between the modules and the flowchart, the flowchart will be explained with the modules as the subject.

[0023] The data refinement module 151 reads sensor data for model learning from the storage area table (FIG. 2) of the sensor data 131 in the storage 130. The type of sensor data for model learning and the range to be read may be based on settings by the user or the like.

[0024] The data refinement module 151 divides the training data into multiple data blocks as preprocessing of the time-series data read as training data (FIG. 5: S500). Methods for dividing the training data include temporal division and classification by clustering. FIG. 6A is a waveform diagram of the sensor time-series data, which is the training data 402. Reference numeral 600 denotes a sliding window, which the data refinement module 151 moves regularly along the time axis, thereby extracting one or more pieces of sensor data within the range enclosed by the window as a single data block. This is an example of a method for dividing the training data over time. The feature quantity of one data block may be, for example, a statistical quantity such as the average value or standard deviation of multiple pieces of data included in the dataset.

[0025] In contrast, clustering is the classification of training data into multiple clusters according to their characteristics. Fig. 6B shows that the training data are classified into four clusters. Of these, the cluster designated by reference numeral 610 contains three data blocks. As shown in Fig. 6C, three pieces of data 610A, 610B, and 610C exist discretely in the training data 402, which is time-series data.

[0026] Figure 7 is a block diagram for explaining the data refining process by data refining module 151, and will be explained in conjunction with the flowchart in Figure 5. As already mentioned, data refining module 151 divides training data 402 into data blocks 7000 of equal size over time (Figure 5: S500).

[0027] Next, the data refinement module 151 narrows down all data blocks 402BL of the learning data to data blocks to be searched for as to whether they are the impact data described above (FIGS. 5 and 7: S502). In FIG. 7, the data blocks indicated by 7001 to 7004 are the search targets. Since checking whether all data blocks of the learning data are impact data would increase the calculation load, only some data blocks are set as search targets. The method for determining the data blocks to be search targets will be described later. Note that this does not prevent all data blocks from being search targets.

[0028] Next, the data refining module 151 selects one dataset 7000 from the datasets to be searched (FIGS. 5 and 7: S504). The model construction module 152 removes the selected data block from the training data (FIGS. 7: S5041), and performs unsupervised machine learning based on the remaining data block (404BL) (FIGS. 5 and 7: S506), to create a failure sign detection model (target model) (FIG. 7: 7010).

[0029] Furthermore, the model construction module 152 performs unsupervised machine learning based on all data blocks (402BL) of the learning data (FIG. 7: S507) to create a failure sign detection model (reference model: FIG. 5: S508, FIG. 7, 7012).

[0030] Next, the controller of the failure prediction server starts the model application module 161, which reads the target model 7010 from the model configuration module 152 and further reads the aforementioned verification data 404 from the storage 130, i.e., time-series data that is sensor data from a predetermined time before the device failure, and applies the target model 7010 to this to calculate the anomaly level (target anomaly level) 7014 of the verification data (FIGS. 5 and 7: S510), which is then stored in memory. Similarly, the model application module 161 applies the reference model 7012 to the verification data 404 to calculate the anomaly level (reference anomaly level) 7016 of the verification data (FIGS. 5 and 7: S512).

[0031] Next, the model application module 161 compares the target anomaly degree 7014 with the reference anomaly degree 7016 and evaluates the difference between them (FIGS. 5 and 7: S514). This evaluation can be performed as follows. 800 is a period (increase period) during which the former increases more than the latter, and 802 is a period (decrease period) during which the former decreases less than the latter. The model application module 161 sets the evaluation index during the increase period to the integral value of the increase period width 800 of the anomaly score increase amount 800A, and sets the evaluation index during the decrease period to the integral value of the decrease period width 802 of the anomaly score decrease amount 802A, and sets the difference between the two to the evaluation result. Note that it is also possible to use the number of data items with an increase in abnormal values ​​and the number of data items with a decrease in abnormal values ​​as the evaluation index, and to use the difference between the two to be the evaluation result.

[0032] In step S514, the model application module 161 finishes evaluating the impact of the selected data block on the degree of anomaly, and then repeats this process until the evaluation of all data blocks is completed (S516), determines the data block with the greatest impact among data blocks 7001 to 7004 as the impact data, and determines the impact of the impact data on the detection accuracy of the failure precursor detection model (S518).

[0033] 9A is a graph comparing the time changes of the anomaly score of the target anomaly degree 7016 and the anomaly score of the reference anomaly degree 7014. The anomaly score of the target anomaly degree 7016 increases at an earlier stage before the day of the equipment failure than the reference anomaly degree 7014, and the anomaly score on the day of the failure also has a larger value. Therefore, the model application module 161 determines that, in order to improve the accuracy of the failure sign detection model, it is more appropriate to construct the failure sign detection model by excluding impact data from the learning data.

[0034] 9B, the target anomaly degree 7016 increases in anomaly score later than the reference anomaly degree 7014 at the stage before the equipment failure day, and the anomaly score on the failure day is also smaller. Therefore, the model application module 161 determines that it is preferable not to remove the impact data from the learning data in order to maintain the accuracy of the failure sign detection model.

[0035] In Figure 9C, there is almost no difference in the degree of change in the anomaly score between the target anomaly level 7016 and the reference anomaly level 7014, including on the day of equipment failure, so the model application module 161 determines that the impact data does not affect the accuracy of the failure sign detection model.

[0036] The data processing system 101 completes step S518 and ends the flowchart in Figure 5. Note that multiple data blocks may be determined as impact data in descending order of impact. The data processing system may start the flowchart in Figure 5 when building a failure sign detection model or when performing maintenance on it.

[0037] 7, it has been explained that the data refining module 151 narrows down the search targets for impact data. Therefore, the data refining module 151 uses the priority of the data blocks 7000 that belong to the learning data 402. The data refining module 151 prioritizes data blocks (7001 to 7004 in FIG. 7) with a predetermined priority or higher when searching for impact data.

[0038] The data refinement module 151 calculates the priority of a data block based on the degree of anomaly (self-anomaly degree) that the data block occupies in the learning data. The higher the self-anomaly degree of a data block, the more likely it is impact data.

[0039] For example, the data refinement module 151 calculates the degree of anomaly of a data block according to the method illustrated in Fig. 10. The data refinement module 151 divides all data blocks 402BL included in the learning data into two groups BL800 and BL802, performs machine learning 810 based on the data block BL800 of one group to create a learning model, and calculates 812 the degree of anomaly of the data block BL802 of the other group.

[0040] The data refinement module 151 calculates the degree of inherent anomaly of each data block for all patterns, although there are multiple combinations of ways to divide all data blocks included in the learning data into two groups. As a result, the degree of inherent anomaly is calculated for all data blocks.

[0041] The data refinement module 151 determines the priority of a data block by summing multiple anomaly degrees for the same data block for all patterns, or by averaging multiple anomaly degrees for the same data block. The model application module 161 searches for impact data on a predetermined number of data blocks in descending order of priority. This reduces the load of determining impact data compared to when all data blocks are searched for impact data.

[0042] When the model application module 161 determines the impact data, it displays it to the user of the data processing system 101. Fig. 11 is an example of a display list as one form of displaying information about impact data. This display example shows that there are four data blocks distinguished by identifiers 1 to 4 as impact data, the accuracy of each piece of impact data, i.e., an index of whether the accuracy of failure sign detection improves or, conversely, worsens when the impact data is deleted from the learning data, and the occurrence rate of the impact data relative to the learning data, i.e., the ratio of the size of the impact data to the size of the learning data.

[0043] FIG. 12 is another example of a screen 1000 of impact data-related information output to a user device, and includes the distribution of impact data for each of sensors 1 and 2. The horizontal axis is the scale of the sensor output value, and the vertical axis is the count number of the output value. 1202 indicates the sensor output, and 1200 indicates the impact data within the sensor output. It can be seen that impact data appears in areas where the sensor output value is large for both sensor 1 and sensor 2. Furthermore, screen 1000 includes the time distribution of the appearance of impact data. The user can learn more about the characteristics and attributes of the impact data based on the display modes of FIGS. 11 and 12.

[0044] By knowing the impact data and the various analytical information related to the impact data described above, the administrative user can easily determine whether it is better to remove the impact data from the learning data and build a failure sign detection model in order to improve the accuracy of the machine learning model's failure sign detection, since the impact data is based on, for example, a malfunction of the device or a change in the device's operating environment, or whether it is better to build a failure sign detection model without excluding the impact data from the learning data, since the impact data accurately represents the state of the device, based on the timing of the appearance of the impact data, etc.

[0045] The present invention is not limited to the above-described embodiments and includes various modifications. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to those including all of the described configurations. Furthermore, it is possible to replace part of the configuration of one embodiment with the configuration of another embodiment, or to add the configuration of another embodiment to the configuration of one embodiment. Furthermore, it is possible to add, delete, or replace part of the configuration of each embodiment with other configurations.

[0046] In the above-described embodiment, the detection of signs of failure in industrial machinery has been described, but the present invention may be applied to detecting signs of abnormality in the vital signs of the human body, not limited to industrial machinery, as long as signs of some abnormality can be detected. Also, in the above-described embodiment, the impact data has been described as consisting of one data block, but the impact data may also be a combination of multiple data blocks. [Explanation of symbols]

[0047] 101: Data processing system, 130: Storage, 150: Data refinement server, 151: Data refinement module, 152: Model construction module, 160: Failure prediction server

Claims

1. receiving time series data from a target measurement sensor; a controller that executes a program to perform data processing for detecting a sign of abnormality in the target based on the time-series data; 1. A data processing system comprising: The controller collecting learning data for machine learning from the received time series data, the learning data including a plurality of measurement data; Performing machine learning based on the learning data to construct a model for detecting signs of abnormality in the target; Applying the model to validation data to calculate the degree of anomaly of the validation data; Identifying measurement data that affects the degree of anomaly from among the measurement data included in the learning data as data that should be noted; Present the relevant noteworthy data, identifying the measurement data as data of interest; selecting at least one measurement data item from the training data; the degree of anomaly resulting from performing the machine learning after excluding the selected measurement data from the learning data is set as a first degree of anomaly, the degree of anomaly resulting from performing the machine learning without excluding the selected measurement data is set as a second degree of anomaly, the first degree of anomaly is compared with the second degree of anomaly, and the data to be noted is determined based on the comparison result; This is done by Furthermore, the controller The verification data is time-series data at the time of a failure as an abnormality of the target. The learning data is collected from time series data of a period that does not include a failure of the target. Data processing system.

2. The controller determining, as the data to be noted, the measurement data that maximizes the difference between the first abnormality degree and the second abnormality degree; 10. The data processing system of claim 1.

3. The controller selecting a plurality of measurement data items from the learning data in a preferential manner; determining the data to be noted from the selected plurality of measurement data; 3. The data processing system of claim 2.

4. The controller The machine learning is performed by unsupervised learning.

10. The data processing system of claim 1.

5. The controller Detecting signs of failure of the target industrial machinery based on time-series data from the industrial machinery; 10. The data processing system of claim 1.

6. The controller Dividing the training data into a plurality of blocks; Identifying the data to be noted from the plurality of blocks; 10. The data processing system of claim 1.

Citation Information

Patent Citations

  • Abnormality detection method and abnormality detection device

    JP2015114967A

  • Abnormality detecting device, abnormality detecting method, and abnormality detecting program

    JP2018112863A

  • Learning data confirmation support apparatus, machine learning apparatus, and failure prediction apparatus

    JP2020102001A

  • Detection device and detection program

    JP2020140580A

  • Projection system, position detection system, and position detection method

    JP2021033554A