Information processing device, drift detection method, and drift detection program

The information processing device uses a future prediction model to anticipate data drift, ensuring timely intervention and maintaining accurate machine learning model performance in plant operations.

JP2025114282APending Publication Date: 2025-08-05YOKOGAWA ELECTRIC CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024008886
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-24
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

Existing data drift detection methods only respond to data drift after it occurs, leading to instability in machine learning model performance, particularly in plant operations, as they rely on detecting outliers in acquired data.

Method used

An information processing device that predicts data drift by using a future prediction model to analyze sensor data from a machine learning model, allowing for advance detection of data drift before it affects model performance.

Benefits of technology

Enables proactive detection of data drift, preventing prolonged operation with inaccurate predictions by alerting operators to retrain or tune the model before performance deterioration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025114282000001_ABST
    Figure 2025114282000001_ABST
Patent Text Reader

Abstract

To achieve advance detection of data drift occurrence.SOLUTION: This information processing device includes: a data acquisition unit configured to acquire sensor data from a sensor that measures the state of a system; and a drift detection unit configured to detect, based on execution results of future predictions performed on the sensor data, whether data drift will occur from a first training dataset used to train a first machine learning model that takes the sensor data as input and outputs a class indicating the state of the system.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, a drift detection method, and a drift detection program. [Background technology]

[0002] From the perspective of implementing MLOps (Machine Learning Operations), the performance of machine learning models is monitored during operation. For example, one of the factors that can cause a deterioration in the performance of a machine learning model is a phenomenon known as data drift, in which the distribution of features in the data input to the machine learning model during testing or system operation changes from the distribution of features in the data used to train the machine learning model.

[0003] One known method for detecting such data drift is outlier detection, which determines whether data acquired during system operation falls outside a set range based on the training dataset of a machine learning model. [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] Yukiko Suga, Satoshi Saeki, Norimitsu Takahashi, Takahiro Tanaka, Yohei Okawa, Kenichi Oguro, Kazuhiro Moriya, Hiroyuki Misaki, Junichiro Kitagawa, Masahisa Moriya, Takuhito Yamanoshita, Naotomo Karibe, Daisuke Takatada, "The Shortest Path to Data Scientist Certification (Literacy Level) Official Reference Book, 2nd Edition," Gijutsu Hyoronsha Publishing, May 27, 2022, 2nd Edition, 1st Printing, p. 138 Summary of the Invention [Problem to be solved by the invention]

[0005] However, in the above-described outlier detection, the occurrence of data drift is detected only when the data corresponding to the outlier is acquired, so there is room for improvement in that the response to data drift must be made after the fact.

[0006] An object of the present invention is to realize advance detection of the occurrence of data drift. [Means for solving the problem]

[0007] An information processing device according to one aspect includes a data acquisition unit that acquires sensor data from a sensor that measures the state of a system, and a drift detection unit that detects, based on the results of performing future predictions on the sensor data, whether data drift has occurred from a first training dataset used to train a first machine learning model that takes the sensor data as input and outputs a class that indicates the state of the system.

[0008] In one aspect of the drift detection method, a computer executes a process of acquiring sensor data from a sensor that measures the state of a system, and detecting, based on the results of performing future predictions on the sensor data, whether data drift has occurred from a first training dataset used to train a first machine learning model that takes the sensor data as input and outputs a class that indicates the state of the system.

[0009] In one aspect, a drift detection program causes a computer to execute a process of acquiring sensor data from a sensor that measures the state of a system, and detecting, based on the results of performing future predictions on the sensor data, whether data drift has occurred from a first training dataset used to train a first machine learning model that takes the sensor data as input and outputs a class that indicates the state of the system. [Effects of the Invention]

[0010] According to one embodiment, it is possible to realize advance detection of the occurrence of data drift. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 1 is a block diagram illustrating an example of a functional configuration of an information processing device. [Figure 2] FIG. 2 is a diagram showing the state of the plant equipment. [Figure 3] FIG. 3 is a diagram showing the trend of the process data. [Figure 4] FIG. 4 is a schematic diagram (1) showing an example of generating a future prediction model. [Figure 5] FIG. 5 is a schematic diagram (2) showing an example of generating a future prediction model. [Figure 6] FIG. 6 is a diagram showing the trend of the process value. [Figure 7] FIG. 7 is a flowchart (1) showing the procedure of the drift detection process. [Figure 8] FIG. 8 is a flowchart (2) showing the procedure of the drift detection process. [Figure 9] FIG. 9 is a diagram (1) showing an example of a threshold value for detecting outliers. [Figure 10] FIG. 10 is a diagram (2) showing an example of a threshold value for detecting outliers. [Figure 11] FIG. 11 is a flowchart showing the procedure of the threshold setting process. [Figure 12] FIG. 12 is a schematic diagram showing an example of generating a new data set. [Figure 13] FIG. 13 is a schematic diagram showing an example of calculation of the amount of change. [Figure 14] FIG. 14 is a flowchart showing the procedure of the outlier detection process. [Figure 15] FIG. 15 is a diagram (3) showing an example of a threshold value for detecting outliers. [Figure 16] FIG. 16 is a diagram illustrating an example of a hardware configuration. DETAILED DESCRIPTION OF THE INVENTION

[0012] Hereinafter, an information processing device, a drift detection method, and a drift detection program according to the present disclosure will be described with reference to the accompanying drawings. Note that this embodiment merely illustrates one example or aspect, and the structure, action, function, properties, characteristics, methods, uses, etc. according to the present disclosure are not limited by such an example.

[0013] <Embodiment 1> <Examples of usage scenarios> Fig. 1 is a block diagram showing an example of the functional configuration of an information processing device. Fig. 1 shows, as an example, an information processing device 10 that provides a drift detection function for detecting data drift that may occur in process data input to a machine learning model (hereinafter referred to as an "anomaly detection model") that detects anomalies in a plant.

[0014] In one embodiment, the information processing device 10 may be realized as a server that provides the drift detection function on-premise. In another embodiment, the information processing device 10 may be realized as a Software as a Service (SaaS) type application, thereby providing a service corresponding to the drift detection function as a cloud service.

[0015] Sensors 20A to 20X may be connected to such an information processing device 10. Hereinafter, when it is not necessary to distinguish between the individual sensors 20A to 20X, the sensors 20A to 20X may be referred to as "sensor 20."

[0016] The information processing device 10 and the sensor 20 may be communicably connected via any network NW. The network NW may be wired or wireless and may be realized by any technology such as Internet technology, industrial communication standards, or low-power wireless communication standards for IoT (Internet of Things).

[0017] The sensor 20 is a measuring device that measures the state of a production process in a plant. For example, the sensor 20 may be realized by an instrumentation device, a so-called field device, that measures temperature, pressure, flow rate, etc. The process data measured by such a sensor 20 may be transmitted to the information processing device 10 at any period, such as 10 seconds, 15 minutes, 1 hour, or 4 hours.

[0018] Here, a plant is used as an example of a system to which the machine learning model is applied, and process data is used as an example of data measuring the state of the system. However, it should be noted that the use scenario of the drift detection function described above is not limited to this. That is, the system to which the machine learning model is applied is not limited to a plant, but may include one or more arbitrary devices or one or more arbitrary pieces of equipment. Furthermore, the data measuring the state of the system is not limited to actual measurements taken by the sensor 20, but may also be predicted values from a simulator or other devices, or index values such as performance or scores.

[0019] <Data Drift> FIG. 2 is a diagram showing the status of equipment in a plant. FIG. 2 shows a graph in which the status of a certain piece of equipment is plotted over time. For example, the horizontal axis of the graph shown in FIG. 2 indicates time, such as date. Furthermore, the vertical axis of the graph shown in FIG. 2 indicates the status of the equipment, and is represented by, for example, a value "1" corresponding to normal and a value "0" corresponding to abnormal.

[0020] Furthermore, FIG. 2 shows the training data period used to train the anomaly detection model and the test data period during which predictions were made by the trained anomaly detection model in a test or production environment, out of the total period during which the sensor 20 measured the process data.

[0021] Here, in FIG. 2, the equipment state output by the anomaly detection model to which the process data measured by the sensor 20 is input is shown by a shaded line, and the actual equipment state is shown by a solid line.

[0022] As shown in Figure 2, it is clear that in the test data period, at a certain point, namely May 21, 2012, the actual equipment status and the equipment status predicted by the anomaly detection model do not match.

[0023] As such, the performance of an anomaly detection model can deteriorate after training is complete. One of the causes of this is data drift. Data drift refers to the phenomenon in which the distribution of features in the data input to a machine learning model during testing or system operation changes from the distribution of features in the data used to train the machine learning model.

[0024] FIG. 3 is a diagram showing the trend of process data. For example, the horizontal axis of the graph shown in FIG. 3 indicates time, such as date, and corresponds to the time series of the graph shown in FIG. 2. Furthermore, the vertical axis of the graph shown in FIG. 3 indicates the value of a process variable, such as a physical quantity such as temperature, pressure, or flow rate. Hereinafter, the value of a process variable may be referred to as a "process value."

[0025] As shown in Figure 3, the point at which predictions by the anomaly detection model no longer match the actual equipment status, i.e., the distribution of process values changes around the date of May 21, 2012, indicating a divergence between the two distributions. Because such events can occur during plant operation, detecting this data drift, which leads to a deterioration in accuracy, is of great technical significance in the operation of an anomaly detection model. Detecting the occurrence of this data drift can prevent plant operation from continuing based on an incorrect prediction result from the anomaly detection model.

[0026] <One aspect of the issue> As explained in the Background Art section above, in the above-mentioned outlier detection, the occurrence of data drift is detected only when the data corresponding to the outlier is acquired, so there is room for improvement in that the response to data drift must be made after the fact.

[0027] Furthermore, data drift is not necessarily detected at the time when an outlier is observed. For example, as the interval at which process data is collected from the sensor 20 becomes longer, the time lag between the time when an outlier is observed in the process data and the time when the above-mentioned outlier detection is performed and data drift is detected increases. Therefore, as the interval at which process data is collected from the sensor 20 becomes longer, the period during which the plant operates in a state where the accuracy of the prediction results by the anomaly detection model is unstable also becomes longer.

[0028] <One aspect of the problem-solving approach> Therefore, the drift detection function according to this embodiment predicts whether or not data drift will occur based on the execution result of future prediction regarding the process data acquired from the sensor 20. This realizes advance detection of the occurrence of data drift.

[0029] <Configuration of information processing device> Next, the functional configuration of the information processing device 10 having the drift detection function will be described. Fig. 1 shows a block diagram related to the drift detection function of the information processing device 10. Note that Fig. 1 only shows an excerpt of functional units related to the drift detection function, and does not preclude the information processing device 10 from being provided with functional units other than those shown.

[0030] As shown in FIG. 1, the information processing device 10 includes a data acquisition unit 11, a memory unit 12, a first training unit 13, a threshold setting unit 14, an outlier detection unit 15, a second training unit 16, a drift detection unit 17, an anomaly detection unit 18, and an output control unit 19.

[0031] The data acquisition unit 11 is a processing unit that acquires process data measured by the sensors 20. In one embodiment, the data acquisition unit 11 can acquire process data for each sensor 20 at any period, such as 10 seconds, 15 minutes, 1 hour, 4 hours, or 1 day. For example, the data acquisition unit 11 acquires process data via a distributed control system such as a DCS (Distributed Control Systems) or a PLC (Programmable Logic Controller). Alternatively, the data acquisition unit 11 can acquire process data via a Plant Information Management System (PIMS) that stores the history of process data.

[0032] As one aspect, the data acquiring unit 11 stores the process data acquired for each sensor 20 in the storage unit 12 as first training data 12A until the learning phase of the anomaly detection model is completed.

[0033] In another aspect, during the inference phase of the above-mentioned anomaly detection model, for example, in a test environment or a production environment, the data acquisition unit 11 stores the process data acquired for each sensor 20 in the memory unit 12 as test data 12B.

[0034] At this time, the data acquiring unit 11 can also calculate feature amounts from the time-series data of the process values and store the calculated feature amounts as the first training data 12A or the test data 12B. For example, the calculation of feature amounts may be realized by pre-processing that quantifies the features of the process data, that is, by so-called feature extraction.

[0035] Such feature extraction can be achieved by statistical processing to calculate statistical values such as mean and variance, frequency conversion to convert time-series data of process values from the time domain to the frequency domain, or convolution using filters or a convolutional layer neural network.

[0036] For example, the data acquiring unit 11 reduces the dimension of the time-series data of the process values by calculating a statistical value, such as a weighted average, from the time-series data of the process values for each of the sensors 20A to 20X for a predetermined period going back from the time the process data was acquired. A feature vector including the weighted average value calculated for each of the sensors 20A to 20X in this manner may be stored as the first training data 12A or the test data 12B.

[0037] Note that, although an example in which one feature amount is calculated for one sensor 20 has been given here, it is of course possible to calculate multiple types of feature amounts for one sensor 20. Also, although an example in which feature amounts for multiple sensors 20 are calculated has been given here, the number of sensors 20 may be one or more, and a usage scenario in which the number of sensors 20 is one is not excluded.

[0038] The first training unit 13 is a processing unit that trains the anomaly detection model. Hereinafter, as an example of the anomaly detection model, a machine learning model that executes a classification task that uses a feature vector obtained from process data as input and outputs a class that classifies the state of a plant, for example, normal or abnormal, will be taken as an example.

[0039] Such an anomaly detection model may be realized by any machine learning model, such as a neural network, a support vector machine, or gradient boosting. For example, a set of first training data 12A, i.e., a first training data set, stored in the storage unit 12 can be used to train the anomaly detection model. For example, the first training data 12A may be data in which a correct label corresponding to a class of the plant state, such as "normal" or "abnormal," is assigned to a feature vector obtained by feature extraction of the process data of each sensor 20.

[0040] Hereinafter, in order to distinguish whether an anomaly detection model has been trained or not, an anomaly detection model before training will be identified as "anomaly detection model 18m" and a trained anomaly detection model will be identified as "anomaly detection model 18M."

[0041] For example, in the learning phase, the feature vectors of the process data included in the first training data 12A are used as explanatory variables of the anomaly detection model 18m, and the correct answer labels are used as objective variables of the anomaly detection model 18m, and the parameters of the anomaly detection model are trained according to an arbitrary machine learning algorithm, for example, deep learning, thereby generating a trained anomaly detection model 18M.

[0042] In the inference phase, the feature vector of the process data included in the test data 12B is input to the anomaly detection model 18M. The anomaly detection model 18M to which the feature vector of the process data has been input outputs a certainty factor for each class of "normal" or "abnormal" indicating the state of the plant. Based on the certainty factor of the "abnormal" state of the plant obtained in this way, the anomaly detection unit 18, which will be described later, determines whether the plant is abnormal. For example, if the certainty factor of the "abnormal" state of the plant is equal to or greater than a threshold, the plant state is determined to be abnormal.

[0043] Although two-class classification into "normal" and "abnormal" classes has been given here as an example of a machine learning task performed by the anomaly detection model, multi-class classification into "steady state," "unsteady state," or "abnormal" may also be performed. Also, although an example has been given here in which the anomaly detection model outputs a certainty factor for each class, it may instead output the class with the highest certainty factor.

[0044] The threshold setting unit 14 is a processing unit that sets a threshold used for the outlier detection. In one embodiment, the threshold setting unit 14 can set upper and lower limit values that define the boundaries of a numerical range that is not an outlier as thresholds for each sensor 20 based on a set of first training data 12A stored in the storage unit 12, i.e., a first training data set. Such thresholds may be set, for example, using a box-and-whisker plot, Hotelling's method, significance level, or the like, which are used as the above-mentioned outlier detection techniques.

[0045] The outlier detection unit 15 is a processing unit that performs the above-described outlier detection. As one aspect, the outlier detection unit 15 performs outlier detection for each sensor 20 on the process value acquired from the sensor 20 during the inference phase of the anomaly detection model 18M, i.e., when the trained anomaly detection model 18M has been generated. For example, the outlier detection unit 15 determines that the process value is an outlier when the process value is below a lower threshold or exceeds an upper threshold. The outlier determination result obtained for each sensor 20 is associated with the feature vector of the process data of each sensor 20 as a ground truth label used for training the machine learning model that performs the above-described future prediction, and is then stored in the storage unit 12 as second training data 12C. Note that, although an example of determining whether a process value is an outlier has been given here, whether a feature value calculated from the process value is an outlier may also be determined.

[0046] The second training unit 16 is a processing unit that trains a machine learning model that performs future predictions regarding process data acquired from the sensor 20. Hereinafter, the machine learning model that performs the above-mentioned future predictions may be referred to as a "future prediction model" in order to distinguish it from the label of the above-mentioned anomaly detection model.

[0047] In one embodiment, the second training unit 16 generates the above-described future prediction model for each sensor 20. Such an anomaly detection model may be realized by any machine learning model, such as a neural network, a support vector machine, or gradient boosting.

[0048] In one aspect, the machine learning task of the future prediction model may be a classification problem. In this case, the future prediction model receives a feature vector of the process data of each sensor 20 as input, and outputs a prediction result of whether an outlier will occur in the process data of a certain sensor 20 after a predetermined period from the time when the process data is measured or acquired, such as whether an outlier exists or not.

[0049] Hereinafter, in order to distinguish whether a future prediction model has been trained or not, a future prediction model before training will be identified as "future prediction model 17mA" and a trained future prediction model will be identified as "future prediction model 17MA."

[0050] FIG. 4 is a schematic diagram (1) showing an example of generating a future prediction model. For example, FIG. 4 shows an example in which the machine learning task of the future prediction model is a classification problem. Furthermore, FIG. 4 also shows a schematic diagram of an example of generating a future prediction model 17MA that performs future predictions on process data of a sensor 20A.

[0051] 4, a set of second training data 12C stored in the storage unit 12, i.e., a second training data set TR21, can be used to train the future prediction model 17mA. Note that, as an example, FIG. 4 illustrates the second training data set TR21 when the interval at which process data is acquired from the sensor 20A is one day.

[0052] For example, the second training data 12C may be data in which the feature vectors of the process data of each sensor 20 are assigned a correct answer label corresponding to the outlier determination result in the process data of sensor 20A a predetermined period of time after the process data was acquired, for example, five days later.

[0053] 4, if the process value of sensor 20A is determined to be an outlier a predetermined period after the process data of each sensor 20 is acquired, a flag "0" indicating this is assigned as a correct label. On the other hand, if the process value of sensor 20A is determined not to be an outlier a predetermined period after the process data of each sensor 20 is acquired, a flag "1" indicating this is assigned as a correct label.

[0054] For example, in the learning phase, the feature vectors of the process data included in the second training data 12C are used as explanatory variables of the future prediction model 17mA, and the correct answer labels are used as objective variables of the future prediction model 17mA, and the parameters of the future prediction model are trained according to an arbitrary machine learning algorithm. As a result, a trained future prediction model 17MA is generated.

[0055] In the inference phase, the feature vector of the process data included in the test data 12B is input to the future prediction model 17MA. The future prediction model 17MA, to which the feature vector of the process data has been input in this way, outputs a prediction result of whether or not an outlier will occur in the sensor 20A after a predetermined period of time, for example, a confidence level for each class of "no outlier" or "outlier present." Based on the confidence level of the prediction result "no outlier" or "outlier present" obtained in this way, the drift detection unit 17, which will be described later, predicts whether or not data drift will occur. For example, if the confidence level of the prediction result "outlier present" is equal to or greater than a threshold, it is determined that data drift will occur after a predetermined period of time.

[0056] Although two-class classification has been given as an example of a machine learning task executed by the future prediction model, multi-class classification may also be executed. Also, although an example has been given in which the future prediction model outputs a confidence level for each class, it may instead output the class with the highest confidence level.

[0057] In another aspect, the machine learning task of the future prediction model may be a regression problem. In this case, the future prediction model receives a feature vector of the process data of each sensor 20 as input and outputs a prediction result, such as a process value, of the process data of the sensor 20 after a predetermined period from the time when the process data is measured or acquired.

[0058] Fig. 5 is a schematic diagram (2) showing an example of generating a future prediction model. For example, Fig. 5 shows an example in which the machine learning task of the future prediction model is a regression problem. Furthermore, Fig. 5, like Fig. 4, also shows a schematic diagram of an example of generating a future prediction model 17MA that performs future prediction on process data of sensor 20A.

[0059] 5, a set of second training data 12C stored in the storage unit 12, i.e., a second training data set TR22, can be used to train the future prediction model 17mA. Note that, as an example, FIG. 5 illustrates the second training data set TR22 when the interval at which process data is acquired from the sensor 20A is one day.

[0060] For example, the second training data 12C may be data in which the feature vector of the process data of each sensor 20 is assigned a correct label corresponding to the value of the process variable of sensor 20A a predetermined period of time after the process data was acquired, for example, five days later.

[0061] In the learning phase, the feature vectors of the process data included in the second training data 12C are used as explanatory variables of the future prediction model 17mA, the correct answer labels are used as objective variables of the future prediction model 17mA, and the parameters of the future prediction model are trained according to an arbitrary machine learning algorithm. As a result, a trained future prediction model 17MA is generated.

[0062] In the inference phase, the feature vector of the process data included in the test data 12B is input to the future prediction model 17MA. The future prediction model 17MA, to which the feature vector of the process data has been input in this manner, outputs a predicted value of the process variable of the sensor 20A after a predetermined period of time. Based on the predicted value of the process variable obtained in this manner, the drift detection unit 17, which will be described later, predicts whether or not data drift will occur. For example, if the predicted value of the process variable exceeds the upper threshold set by the threshold setting unit 14, or if the predicted value of the process variable is less than the lower threshold set by the threshold setting unit 14, it is determined that data drift will occur after the predetermined period of time.

[0063] Although the example in which the future prediction model outputs a predicted value of a process variable has been described above, the present invention is not limited to this. For example, the future prediction model may also output a predicted value of a feature quantity calculated from a process value.

[0064] The drift detection unit 17 is a processing unit that predicts the occurrence of data drift based on the execution result of future prediction of process data for each sensor 20. In one aspect, the drift detection unit 17 executes the following determination for each sensor 20A to 20X. For example, the drift detection unit 17 inputs the feature vector of the process data acquired for each sensor 20 by the data acquisition unit 11 to the future prediction model 17M.

[0065] In one aspect, when the machine learning task of the above-described future prediction model is a classification problem, the future prediction model 17M outputs a prediction result of whether or not an outlier will occur in the sensor 20 after a predetermined period of time, for example, a confidence level for each class of "no outlier" or "outlier present." In this case, the drift detection unit 17 determines that data drift will occur after the predetermined period of time if the confidence level of the prediction result of "outlier present" or "no outlier present" of whether or not an outlier will occur in the sensor 20 after a predetermined period of time is equal to or greater than a threshold.

[0066] In another aspect, when the machine learning task of the above-mentioned future prediction model is a regression problem, the future prediction model 17M outputs a predicted value of the process variable of the sensor 20 after a predetermined period of time. At this time, if the predicted value of the process variable of the sensor 20 after the predetermined period of time exceeds the upper threshold set by the threshold setting unit 14, or if the predicted value of the process variable is less than the lower threshold set by the threshold setting unit 14, the drift detection unit 17 determines that data drift will occur after the predetermined period of time.

[0067] Such a determination is performed for each of the X future prediction models 17MA to 17MX corresponding to the sensors 20A to 20X. At this time, if the occurrence of data drift after a predetermined period is detected for at least one or more sensors 20, an alert is output by the output control unit 19, which will be described later.

[0068] The anomaly detection unit 18 is a processing unit that detects anomalies in the plant. As one aspect, in the inference phase of the anomaly detection model 18M, for example, in a test environment or a production environment, the anomaly detection unit 18 inputs a feature vector of the process data acquired for each sensor 20 by the data acquisition unit 11 to the anomaly detection model 18M. The anomaly detection model 18M to which the feature vector of the process data has been input in this manner outputs a certainty factor for each class, "normal" or "abnormal," indicating the state of the plant. At this time, the anomaly detection unit 18 determines that the state of the plant is abnormal if the certainty factor for the "abnormal" state of the plant is equal to or greater than a threshold.

[0069] The output control unit 19 is a processing unit that controls the output of various types of information. The "output" referred to here is not limited to display output, but may also be audio output or print output. Furthermore, the destination of the information output is not limited to a peripheral device connected to the information processing device 10, but may also be an external terminal device, a back-end service, or an application.

[0070] As one aspect, the output control unit 19 can output an alert to any output destination when an outlier is detected by the outlier detection unit 15. Such an alert may include, for example, not only a warning of the occurrence of data drift, but also information that an outlier in the actual measurement value has been detected, the acquisition time of the process value at which the outlier was detected, the type of process variable at which the outlier was detected, and the like.

[0071] In another aspect, the output control unit 19 can output an alert to any output destination when the drift detection unit 17 detects the occurrence of data drift after a predetermined period of time in at least one sensor 20. Such an alert may include, for example, not only a warning of the occurrence of data drift, but also information that an outlier of the predicted value has been detected, the predicted time at which the outlier will be detected, the type of process variable in which the outlier has been detected, and the like.

[0072] Fig. 6 is a diagram showing the trend of a process value. Fig. 6 shows a graph in which the trends of the actual measured value and predicted value of a process variable corresponding to a certain sensor 20 are plotted using solid and dashed lines. Fig. 6 also shows the upper and lower thresholds for detecting outliers, which are compared with the process value of a certain sensor 20, using dashed and dotted lines. Fig. 6 shows an example in which the interval at which process data is acquired from the sensor 20 is one day.

[0073] Here, Fig. 6 shows an example in which the time when the latest process value was acquired, i.e., the time when the latest future prediction was performed, is February 11, 2012. In this case, the latest actual measured values of the process variable are plotted up to February 11, 2012, while the latest predicted values of the process variable are plotted up to February 16, 2012.

[0074] As shown in Figure 6, as of February 11, 2012, when the latest process value was acquired, the latest measured value of the process variable did not exceed the upper threshold nor fall below the lower threshold, so it was not detected as an outlier.

[0075] On the other hand, the latest predicted value of the process variable exceeds the upper threshold on February 16, 2012, five days after February 11, 2012. Therefore, the occurrence of data drift in the feature corresponding to the process variable of sensor 20 can be detected in advance as of February 11, 2012.

[0076] This makes it possible to start responding to data drift before it actually occurs. For example, by displaying an alert on the display unit of the DCS, workers can check the on-site status of the process where data drift is occurring. In addition, by displaying an alert on the display unit of the HMI (Human Machine Interface) used by process engineers and operators, it is possible to prompt them to retrain or tune the anomaly detection model.

[0077] The output destination of these alerts can be changed depending on the type of feature for which data drift is detected. For example, the alerts can be output to people involved in the process corresponding to the feature for which data drift is detected, such as process engineers, operators, or workers.

[0078] <Processing flow> Next, the flow of processing by the information processing device 10 according to this embodiment will be described. Figures 7 and 8 are flowcharts (1) and (2) showing the procedure of the drift detection process. As shown in Figure 7, the data acquisition unit 11 acquires process data from each sensor 20 (step S101).

[0079] At this time, if the trained anomaly detection model 18M has not been generated (No in step S102), the data acquiring unit 11 executes the following process: The data acquiring unit 11 stores the first training data 12A, which is generated by performing feature extraction from the process data acquired for each sensor 20 in step S101, in the storage unit 12 (step S103).

[0080] Thereafter, if the number of samples of the first training data reaches a specified value (step S104 Yes), the first training unit 13 generates a trained anomaly detection model 18M using the first training data 12A stored in the memory unit 12 (step S105).

[0081] Next, the threshold setting unit 14 sets upper and lower limit values that define the boundaries of a numerical range that is not an outlier as thresholds for each sensor 20 based on the set of first training data 12A, i.e., the first training data set (step S106).

[0082] If the number of samples of the first training data has not reached the specified value (No in step S104), the process proceeds to step S101 without executing the processes in steps S105 and S106.

[0083] Furthermore, if the trained anomaly detection model 18M has already been generated (Yes in step S102), the data acquiring unit 11 executes the following process: That is, as shown in FIG. 8, the data acquiring unit 11 stores test data generated by extracting features from the process data acquired for each sensor 20 in step S101 in the storage unit 12 (step S107).

[0084] Next, the outlier detection unit 15 performs outlier detection for the process values acquired from each sensor 20 (step S108). After that, the outlier detection unit 15 associates the determination result of the outlier obtained for each sensor 20 with the feature vector of the process data of each sensor 20 as a correct label to be used for training the future prediction model, and stores the result as second training data 12C in the storage unit 12 (step S109).

[0085] If the outlier detection results for all sensors 20 are not OK, i.e., if there is a sensor 20 in which an outlier has been detected (step S110 No), the output control unit 19 outputs an alert to warn of the occurrence of data drift (step S111) and terminates the processing.

[0086] On the other hand, if the outlier detection results for all sensors 20 are OK, that is, if there is no sensor 20 in which an outlier has been detected (Yes in step S110), the process branches depending on whether the future prediction model 17M has already been generated.

[0087] At this time, if a future prediction model has not yet been generated (step S112 No), the second training unit 16 determines whether the number of samples of the second training data accumulated in the memory unit 12 has reached a specified value (step S113).

[0088] Then, if the number of samples of the second training data has reached a specified value (Yes in step S113), the second training unit 16 uses the second training data 12C stored in the memory unit 12 to generate a trained future prediction model 17M (step S114).

[0089] If a future prediction model has already been generated (Yes in step S112), the process skips steps S113 and S114 and proceeds to step S115.

[0090] Thereafter, the drift detection unit 17 inputs the feature vector of the test data stored in step S107 into each of the X future prediction models 17M corresponding to the sensors 20A to 20X, thereby performing future prediction of the process data for each sensor 20 (step S115).

[0091] Here, if data drift is detected in the future prediction of the process data for any one of all the sensors 20 (No in step S116), the output control unit 19 outputs an alert to warn of the occurrence of future data drift (step S117).

[0092] Furthermore, if no data drift is detected in the future prediction of the process data for all the sensors 20 (Yes in step S116), the process skips step S117 and proceeds to step S118.

[0093] Thereafter, the anomaly detection unit 18 inputs the feature vector of the test data stored in step S107 into the anomaly detection model 18M, thereby executing anomaly detection for the plant (step S118), and returns to the processing of step S101.

[0094] <One aspect of the effect> As described above, the information processing device 10 according to this embodiment predicts whether or not data drift will occur based on the execution result of future prediction regarding process data acquired from the sensor 20. Therefore, the information processing device 10 according to this embodiment can realize advance detection of the occurrence of data drift.

[0095] <Embodiment 2> Although the embodiments of the present disclosure have been described above, various applications are possible, and further, the present disclosure may be implemented in various different forms other than the above-described embodiments.

[0096] <Example of application of threshold for outlier detection> The second training data 12C can be used not only to generate the future prediction model described above, but also to optimize the threshold setting for outlier detection.

[0097] As described in the first embodiment above, the threshold for outlier detection can be set using a box plot, Hotelling's method, or other methods used to detect outliers. However, the box plot and Hotelling's method do not necessarily allow for setting an optimal threshold.

[0098] That is, when a machine learning model is implemented, it is not always possible to prepare a sufficient number of samples of the first training data 12A. In such cases, there is room for improvement in the box plot and the Hotelling's method, as will be explained below.

[0099] 9 and 10 are diagrams (1) and (2) showing examples of thresholds for outlier detection. For example, FIGS. 9 and 10 show graphs in which each sample of the first training data 12A is plotted. Furthermore, the graphs shown in FIGS. 9 and 10 show a range of values that are not identified as outliers, indicated by hatching. Note that the horizontal axis of the graphs shown in FIGS. 9 and 10 indicates time, for example, date. Furthermore, the vertical axis of the graphs shown in FIGS. 9 and 10 indicates a process value, for example, a physical quantity such as temperature, pressure, or flow rate.

[0100] Here, Fig. 9 shows the threshold value set based on the box plot. This method is an outlier detection method using the difference between the first and third quartiles (IQR). Therefore, if it is not possible to prepare much data from the tail of the distribution of samples in the first training data 12A, as shown in Fig. 9, even a part of the first training data 12A will be subject to outlier detection, which leaves room for improvement.

[0101] FIG. 10 also shows a threshold value set based on the Hotelling method. This method assumes that the data set follows a normal distribution. Therefore, if the distribution of the samples of the first training data 12A follows a normal distribution, as shown in FIG. 10, the maximum or minimum value of the first training data included in the first training data set may approximately coincide with the upper or lower limit of the outlier detection threshold. In other words, if a value that exceeds the maximum or minimum value even slightly is obtained, it will be determined that data drift has occurred. Therefore, since the number of training data samples is generally limited, it is difficult to ensure that the samples of the first training data 12A have a distribution that covers the entire population. Therefore, there is room for improvement in threshold setting based on the Hotelling method.

[0102] From these facts, the threshold setting unit 14 sets a threshold for outlier detection based on the amount of change in each piece of second training data 12C relative to the average value of the process values of the entire second training data 12C.

[0103] Fig. 11 is a flowchart showing the steps of the threshold setting process. As an example, this process can be executed as the process of step S106 in the flowchart shown in Fig. 7. As shown in Fig. 11, if the machine learning task of the future prediction model is a classification problem (Yes in step S201), the threshold setting unit 14 classifies the second training data 12C stored in the storage unit 12 into datasets for each of K classes (step S202).

[0104] Next, the threshold setting unit 14 executes loop processing 1, which repeats the processing of step S203 below a number of times corresponding to the number K of classes classified in step S202. That is, the threshold setting unit 14 excludes data near the center of the data set corresponding to the k-th class (step S203).

[0105] Here, if the machine learning task of the future prediction model is a regression problem (step S201 No), the second training data 12C is not classified and is therefore treated as one data set, that is, one class (K=1).

[0106] By repeating this loop process 1, data near the center is excluded for each data set corresponding to the K classes. Note that the process of step S203 above does not necessarily have to be repeated, and can be executed in parallel for each of the K classes.

[0107] Thereafter, the threshold setting unit 14 generates a new data set by integrating the remaining second training data that have not been excluded from the data sets corresponding to the K classes (step S204).Then, the threshold setting unit 14 calculates the average Ave. value of the process values of the entire new data set generated in step S204. train is calculated (step S205).

[0108] Fig. 12 is a schematic diagram showing an example of generating a new dataset. Fig. 12 shows an example of a second training dataset in which the number K of classes of correct labels used for training by the future prediction model is either "OK (no outliers)" or "NG (outliers)."

[0109] As shown in FIG. 12, in step S202, each of the second training data included in the original second training data set is classified into a data set of class "OK" and a data set of class "NG." Next, in step S203, the second training data for each of the data sets of class "OK" and class "NG" is sorted in descending order of process value, and a predetermined percentage of the data set, for example, 20%, of samples is excluded, centered on a statistical value, for example, the median. Then, in step S204, the remaining second training data from each of the data sets of class "OK" and class "NG" that have not been excluded are integrated into a single data set, thereby generating a new data set. Then, in step S205, the second training data included in the new data set is used as a population to calculate the average process value Ave. train is calculated.

[0110] Returning to the explanation of Figure 11, the threshold setting unit 14 executes loop processing 2, which repeats the following step S206 and the following step S207 a number of times corresponding to the number M of second training data included in the new dataset generated in step S205.

[0111] That is, the threshold setting unit 14 excludes the i-th second training data from the new data set generated in step S205 (step S206). Then, the threshold setting unit 14 determines the average value Ave. of the process values of the entire new data set by excluding the i-th second training data. train The amount of change Δ i is calculated (step S207).

[0112] Figure 13 shows the change Δ i13 is a schematic diagram showing an example of calculation of the average value Ave. of the process values of the entire new data set. As shown in Fig. 13, in step S206, the second training data i=1 is excluded from the M pieces of second training data included in the new data set. Then, in step S207, the second training data i=1 is excluded, and the average value Ave. of the process values of the entire new data set is calculated. train The process of steps S206 and S207 is repeated until the loop counter i reaches M. By repeating this loop process 2, the changes Δ1 to Δ M is calculated.

[0113] Returning to the explanation of FIG. 11, the threshold setting unit 14 calculates the variations Δ1 to Δ2 of the M pieces of second training data obtained in the loop process 2. M Average value of Ave. Δi and standard deviation σ Δi is calculated (step S208).

[0114] Then, the threshold setting unit 14 sets the average value Ave. calculated in step S208. Δi and standard deviation σ Δi The threshold value Th for detecting outliers is calculated based on the above (step S209), and the process ends.

[0115] For example, as an example of the threshold Th for outlier detection, Ave. Δi ±4σ Δi Here, the average value Ave. can be set as the threshold value Th for outlier detection. Δi and standard deviation σ Δi Although an example using both of these has been given, it is also possible to use only one of them.

[0116] Using the outlier detection threshold value Th set in this manner, it is possible to perform outlier detection in step S108 shown in Fig. 8 and drift detection in step S116 shown in Fig. 8. The outlier detection and drift detection are common except for whether the value compared with the outlier detection threshold value Th is an actual measured value of the process value or a predicted value of the process value.

[0117] Fig. 14 is a flowchart showing the procedure for outlier detection processing. As an example, this processing can be performed in outlier detection in step S108 shown in Fig. 8 and in drift detection in step S116 shown in Fig. 8, using the outlier detection threshold Th set according to the procedure in the flowchart shown in Fig. 11. These outlier detection and drift detection are common except for whether the value compared with the outlier detection threshold Th is an actual measured value of the process variable or a predicted value of the process variable.

[0118] As shown in FIG. 14, data to be judged, for example, actual measured values of process variables in test data or predicted values of process variables output by a future prediction model, are added to a new data set (step S301).

[0119] Next, in step S301, the data to be judged is added, and the average value Ave. of the process values of the entire new data set is calculated. train The amount of change Δ j is calculated (step S302).

[0120] Here, the change amount Δ j is the threshold value Th for outlier detection, i.e., Ave. Δi ±4σ Δi If the change amount Δ j is the threshold value Th for outlier detection, i.e., Ave. Δi ±4σ ΔiIf it deviates from the above (No in step S303), it is determined to be an outlier (step S305).

[0121] FIG. 15 is a diagram (3) showing an example of a threshold value for outlier detection. For example, FIG. 15 shows a graph in which each sample of the second training data 12C is plotted. Furthermore, the graph shown in FIG. 15 shows, by hatching, a numerical range that is identified as not being an outlier according to the threshold value Th for outlier detection. Note that the horizontal axis of the graph shown in FIG. 15 indicates time, for example, date. Also, the vertical axis of the graph shown in FIG. 15 indicates a process value, for example, a physical quantity such as temperature, pressure, or flow rate.

[0122] Fig. 15 shows this application example, i.e., the threshold value Th for outlier detection that is set according to the procedure of the flowchart shown in Fig. 11. As shown in Fig. 15, this application example allows the threshold value Th to be set with a margin for the distribution of the second training data, and furthermore, the threshold value Th is not significantly affected even if the number of points in the second training data is very small. In other words, this application example can be used even when the number of samples in the second training data is small and the distribution of the samples in the second training data is not comprehensive.

[0123] <Numbers, etc.> The matters described in the above embodiment, such as the number of sensors 20, the types of anomaly detection models and future prediction models, and specific examples of machine learning tasks, are merely examples and can be changed. Also, the order of processing in the flowcharts described in the embodiment can be changed within a consistent range.

[0124] <System> The information including the processing procedures, control procedures, specific names, various data, and parameters shown in the above documents and drawings can be changed as desired unless otherwise specified. For example, any one or more of the functional units of the information processing device 10, including the data acquisition unit 11, the first training unit 13, the threshold setting unit 14, the outlier detection unit 15, the second training unit 16, the drift detection unit 17, the anomaly detection unit 18, and the output control unit 19, may be configured as separate devices.

[0125] Furthermore, the components of each device shown in the figure are functional concepts and do not necessarily have to be physically configured as shown. In other words, the specific form of distribution and integration of each device is not limited to that shown. In other words, all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc. Note that each configuration may also be a physical configuration.

[0126] Furthermore, each processing function performed by each device can be realized, in whole or in part, by a CPU (Central Processing Unit) and a program analyzed and executed by the CPU, or can be realized as hardware using wired logic.

[0127] <Hardware> Next, an example of the hardware configuration of the computer described in the above embodiment will be described. Fig. 16 is a diagram showing an example of the hardware configuration. As shown in Fig. 16, an information processing device 10 has a communication device 10a, a storage device 10b, a memory 10c, and a processor 10d. Note that the components shown in Fig. 16 may be connected to each other via a bus or the like.

[0128] The communication device 10a is a network interface card or the like. The storage device 10b is a storage device such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive). This storage device does not have to be an internal storage of the information processing device 10, but may be an external or auxiliary storage. For example, the storage device 10b stores programs and DBs that execute the processes shown in FIGS. 7, 8, 11, and 14.

[0129] The processor 10d is a hardware processor that reads a program that executes the same processing as the processing unit shown in Fig. 1 from the storage device 10b or the like and loads it into the memory 10c, thereby operating a process that executes the functions described in Fig. 1.

[0130] Such a process realizes the same functions as the processing units of the information processing device 10. For example, the processor 10d reads from the storage device 10b or the like a program having the same functions as the data acquisition unit 11, the first training unit 13, the threshold setting unit 14, the outlier detection unit 15, the second training unit 16, the drift detection unit 17, the anomaly detection unit 18, the output control unit 19, etc. Then, the processor 10d executes a process that executes the same processing as the data acquisition unit 11, the first training unit 13, the threshold setting unit 14, the outlier detection unit 15, the second training unit 16, the drift detection unit 17, the anomaly detection unit 18, the output control unit 19, etc.

[0131] In this way, the information processing device 10 operates as an information processing device that executes a calculation method by reading and executing a program. The information processing device 10 can also realize functions similar to those of the above-described embodiment by reading the program from a recording medium using a medium reading device and executing the read program. Note that the program in these other embodiments is not limited to being executed by the information processing device 10. For example, the present invention can also be applied in the same way to cases where another computer or server executes the program, or where these execute the program in cooperation with each other.

[0132] The above program can be distributed via a network such as the Internet. The above program can also be recorded on any recording medium and executed by a computer by reading it from the recording medium. For example, the recording medium can be a hard disk, a flexible disk (FD), a CD-ROM, a magneto-optical disk (MO), a digital versatile disk (DVD), or the like.

[0133] <Other> Some examples of combinations of the disclosed technical features are set out below.

[0134] (1) a data acquisition unit that acquires sensor data from a sensor that measures the state of the system; a drift detection unit that detects, based on a result of executing a future prediction on the sensor data, whether or not data drift occurs from a first training dataset used to train a first machine learning model that takes sensor data as input and outputs a class that indicates the state of the system; and An information processing device comprising:

[0135] (2) The information processing device described in (1) is characterized in that the drift detection unit inputs the sensor data acquired by the data acquisition unit into a second machine learning model that takes sensor data as input and outputs a class indicating whether an outlier will occur after a predetermined period of time after the sensor data is acquired, and detects whether the data drift will occur based on the class output by the second machine learning model.

[0136] (3) The information processing device described in (2) is characterized in that it further has a training unit that trains the parameters of the second machine learning model by performing machine learning in which the feature obtained from the sensor data is used as an explanatory variable of the second machine learning model and the correct label of the class indicating whether or not an outlier occurs after a predetermined period of time after the sensor data is acquired is used as the objective variable of the second machine learning model.

[0137] (4) The information processing device described in (1) is characterized in that the drift detection unit inputs the sensor data acquired by the data acquisition unit to a second machine learning model that inputs sensor data and outputs sensor data a predetermined period after the sensor data was acquired, and detects whether the data drift has occurred based on the value of the sensor data output by the second machine learning model after the predetermined period.

[0138] (5) The information processing device described in (4) is characterized in that it further has a training unit that trains the parameters of the second machine learning model by performing machine learning in which the feature obtained from the sensor data is used as an explanatory variable of the second machine learning model and the correct label of the sensor data after a predetermined period after the sensor data is acquired is used as the objective variable of the second machine learning model.

[0139] (6) The information processing device described in (5) is characterized in that, when each of the sensor data included in the second training data set is excluded from the entire second training data set used to train the second machine learning model, the information processing device further has a threshold setting unit that sets a threshold to be compared with the value of the sensor data output by the second machine learning model after a predetermined period based on the amount of change from the average value of the sensor data of the entire second training data set.

[0140] (7) The information processing device described in (6) is characterized in that the threshold setting unit excludes a predetermined percentage of sensor data from the second training data set, centered around the median or mean value, and then calculates the amount of change.

[0141] (8) The information processing device according to any one of (1) to (7), further comprising an output control unit that outputs an alert when the occurrence of the data drift is detected.

[0142] (9) Acquire sensor data from sensors that measure the state of the system; Based on the execution result of the future prediction on the sensor data, detect whether or not data drift occurs from a first training dataset used to train a first machine learning model that takes the sensor data as input and outputs a class that indicates the state of the system. A drift detection method characterized in that the processing is executed by a computer.

[0143] (10) Acquire sensor data from sensors that measure the state of the system; Based on the execution result of the future prediction on the sensor data, detect whether or not data drift occurs from a first training dataset used to train a first machine learning model that takes the sensor data as input and outputs a class that indicates the state of the system. A drift detection program that causes a computer to execute processing. [Explanation of symbols]

[0144] 10. Information processing equipment 11 Data Acquisition Section 12 Storage section 12A First training data 12B test data 12C Second training data 13 First Training Division 14 Threshold setting section 15 Outlier detection section 16 Second Training Division 17 Drift detection unit 17M Future Prediction Model 18 Abnormality detection unit 18M Anomaly Detection Model 19 Output control section 20 sensors

Claims

1. a data acquisition unit that acquires sensor data from a sensor that measures the state of the system; a drift detection unit that detects, based on a result of executing a future prediction on the sensor data, whether or not data drift occurs from a first training dataset used to train a first machine learning model that takes sensor data as input and outputs a class that indicates the state of the system; and An information processing device comprising:

2. The information processing device according to claim 1, characterized in that the drift detection unit inputs the sensor data acquired by the data acquisition unit to a second machine learning model that takes sensor data as input and outputs a class indicating whether an outlier will occur after a predetermined period of time after the sensor data is acquired, and detects whether the data drift will occur based on the class output by the second machine learning model.

3. The information processing device according to claim 2, further comprising a training unit that trains parameters of the second machine learning model by performing machine learning using features obtained from the sensor data as explanatory variables of the second machine learning model and a correct answer label of a class indicating whether or not an outlier occurs after a predetermined period of time after the sensor data is acquired as a target variable of the second machine learning model.

4. The information processing device according to claim 1, characterized in that the drift detection unit inputs the sensor data acquired by the data acquisition unit into a second machine learning model that inputs sensor data and outputs sensor data a predetermined period after the sensor data was acquired, and detects whether the data drift has occurred based on the value of the sensor data output by the second machine learning model after the predetermined period.

5. The information processing device according to claim 4, further comprising a training unit that trains parameters of the second machine learning model by performing machine learning in which the feature quantities obtained from the sensor data are used as explanatory variables of the second machine learning model and the correct label of the sensor data a predetermined period after the sensor data is acquired is used as the objective variable of the second machine learning model.

6. The information processing device according to claim 5, further comprising a threshold setting unit that sets a threshold to be compared with the value of the sensor data after a predetermined period output by the second machine learning model based on an amount of change from the average value of the sensor data of the entire second training data set when each piece of sensor data included in the second training data set is excluded from the entire second training data set used to train the second machine learning model.

7. The information processing device according to claim 6 , wherein the threshold setting unit excludes a predetermined percentage of sensor data from the second training data set centered around a median or an average value, and then calculates the amount of change.

8. The information processing apparatus according to claim 1 , further comprising an output control unit that outputs an alert when the occurrence of the data drift is detected.

9. Acquires sensor data from sensors that measure the state of the system, Based on the results of executing the future prediction on the sensor data, detect whether or not data drift occurs from a first training dataset used to train a first machine learning model that takes the sensor data as input and outputs a class that indicates the state of the system. A drift detection method characterized in that the processing is executed by a computer.

10. Acquires sensor data from sensors that measure the state of the system, Based on the results of executing the future prediction on the sensor data, detect whether or not data drift occurs from a first training dataset used to train a first machine learning model that takes the sensor data as input and outputs a class that indicates the state of the system. A drift detection program that causes a computer to execute a process.