Analysis apparatus, analysis method, and program
The analysis device automates the evaluation of relationships among multiple parameters in large-scale plants, addressing inefficiencies in existing methods by preprocessing and network analysis to enhance accuracy and identify hidden connections.
Patent Information
- Application Number
- JP2023190490
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-08
- Publication Date
- 2025-05-20
AI Technical Summary
Existing methods for analyzing relationships among a large number of parameters in large-scale plants are inefficient, requiring significant manual effort and prone to missing important relationships, especially between parameters not directly connected or having time lags, leading to reduced accuracy and increased computational burden.
An analysis device and method that includes data acquisition, preprocessing, and network analysis to automatically evaluate relationships among multiple parameters, handling time lags and non-stationary data, and outputting network diagrams to visualize these relationships.
Enables efficient analysis of relationships among a large number of parameters without manual selection, identifying previously unnoticed connections and improving accuracy by preprocessing to handle time lags and non-stationary data, thus facilitating better plant operation and preventing equipment damage.
Smart Images

Figure 2025078140000001_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to an analysis device, an analysis method, and a program. [Background technology]
[0002] In operating a large-scale plant, it is important to analyze causal and correlational relationships between parameters such as temperature, pressure, and flow rate, and to utilize the results for plant operation. For example, if there is a causal relationship between parameter 1 and parameter 2, and the value of parameter 2 deviates from the normal range, it is very important to focus on parameter 1 to identify the cause of the abnormality, and to perform appropriate operations so that parameters 1 and 2 are both within the normal range, thereby preventing equipment damage and avoiding operation shutdowns. In general, in large-scale plants, time-series data for a huge number of parameters (more than 100 items) is obtained from measuring instruments, etc. Operators evaluate the causal and correlational relationships between the data based on past performance and knowledge, select parameters according to the plant's condition, and perform operations on the selected parameters. This allows appropriate control of not only the parameter in question but also other related parameters, thereby achieving stable plant operation.
[0003] In such a method, the parameters related to other equipment connected to the equipment related to the selected parameter are often focused on as other parameters having a relationship. Conversely, it is difficult to focus on parameters related to equipment not connected to the equipment related to the selected parameter in the system and consider their influence. In response to this, for example, Patent Document 1 discloses a technique for evaluating the relationship between parameters by calculating the correlation coefficient between the time series data of the parameter of interest and the time series data of other parameters by machine learning or the like and extracting the correlated parameters. By using such a technique, it is possible to analyze the relationship between parameters related to equipment not connected to the system. However, since the parameters to be analyzed for the relationship must be selected by a person, in a large-scale plant that handles a huge number of parameters, it takes a huge amount of work to select parameters, and there is a possibility that selection will be missed due to human work. In addition, a parameter located upstream of the system that has no direct relationship with the target parameter may act on the target parameter via other parameters, but in the conventional method, it is difficult to analyze the relationship with such parameters that do not have a direct correlation. In addition, although it is technically possible to perform machine learning or the like on a huge number of parameters without prior preprocessing, there is a possibility that the calculation time will be huge and the accuracy of the analysis results will be reduced. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Patent Publication No. 2022-74890 Summary of the Invention [Problem to be solved by the invention]
[0005] There is a need for a method for analyzing relationships among a large number of parameters without having to select the parameters for which relationships are to be analyzed.
[0006] The present disclosure provides an analysis device, an analysis method, and a program that can solve the above problems. [Means for solving the problem]
[0007] The analysis apparatus of the present disclosure includes a data acquisition unit that acquires data for each of a plurality of parameters for which relationships are to be analyzed; a pre-processing unit that determines whether relationship analysis is possible for the data and, for the data that requires pre-processing for relationship analysis, performs pre-processing so that relationship analysis is possible; an analysis unit that analyzes the relationships of the data that is determined to be capable of analysis and the data that has been pre-processed; and an output unit that outputs information indicating the relationships of the parameters based on the results of the analysis.
[0008] The analysis method disclosed herein is an analysis method executed by a computer, which obtains data of multiple parameters for which relationships are to be analyzed, determines whether analysis of the relationships is possible for the data, performs preprocessing on the data that requires preprocessing for the relationship analysis so that the relationship can be analyzed, analyzes the relationships between the data determined to be capable of analysis and the data that has been preprocessed, and outputs information indicating the relationships between the parameters based on the results of the analysis.
[0009] The program disclosed herein causes a computer to function as: a means for acquiring data of multiple parameters for which relationships are to be analyzed; a means for determining whether relationship analysis is possible for the data, and for the data that requires preprocessing for relationship analysis, performing preprocessing so that relationship analysis is possible; a means for analyzing the relationships between the data determined to be capable of analysis and the data that has been preprocessed; and a means for outputting information indicating the relationships between the parameters based on the results of the analysis. Effect of the Invention
[0010] According to the above-mentioned analysis device, analysis method, and program, it is possible to analyze relationships among a large amount of data without selecting data for which relationships are to be analyzed. It is also possible to analyze relationships between the target data and data of devices, etc., that are located upstream or downstream of the system and that do not have a direct relationship with the target data. [Brief description of the drawings]
[0011] [Figure 1] FIG. 2 is a block diagram showing an example of an analysis device according to an embodiment. [Diagram 2] FIG. 2 is a diagram illustrating an example of a plant that is a target of analysis processing according to the embodiment. [Diagram 3] FIG. 13 is a diagram illustrating an example of a network diagram obtained by the analysis process of the embodiment. [Figure 4] FIG. 13 is a diagram illustrating an example of a pie chart generated by the analysis process of the embodiment. [Diagram 5] 11 is a flowchart illustrating an example of a relationship analysis process according to the embodiment. [Figure 6] 11 is a flowchart illustrating an example of pre-processing in the analysis process according to the embodiment. [Figure 7] FIG. 2 is a diagram illustrating an example of a hardware configuration of an analysis apparatus according to an embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0012] <Embodiment> The analysis device according to the present disclosure will now be described with reference to FIGS. (composition) FIG. 1 is a block diagram showing an example of an analysis device according to a first embodiment. The analysis device 10 acquires measured values of more than several thousand parameters (temperature, pressure, flow rate, etc.) measured in a large-scale plant, and analyzes relationships such as causal relationships and correlations between these parameters. In the analysis, the user does not need to select the parameters to be analyzed for relationships, and the relationship between parameters related to each of multiple devices that are not connected in the plant system can be automatically analyzed. FIG. 2 shows a schematic configuration of a part of a large-scale plant. The plant 100 includes devices 1 to 7. The devices 1 and 2 to 4 are directly connected by piping, and the devices 2 to 4 and the devices 5 and 6 are also directly connected. For example, a fluid generated by the device 1 is supplied to the devices 2 to 4, processed by the devices 2 to 4, and supplied to the device 5. The devices 2 to 4 are upstream of the device 5, and the device 1 is upstream of the devices 2 to 4. The device 7 is the most upstream device of the device 5, and the device 6 is an upstream device that is directly related to the device 5. The plant 100 has many other devices (not shown), and the plant 100 is operated by monitoring and controlling a huge number of parameters. For example, the plant 100 has several thousand or more parameters, and each parameter is sampled, for example, at one-second intervals and transmitted to a control device. In this embodiment, a network diagram showing the relationship between parameters, taking into account the parameters of devices not connected in the system, is created from the correlation of data (time-series data) between these several thousand or more parameters, and it is possible to perform a causal analysis of events occurring in the plant 100. An example of the network diagram is shown in FIG. 3. p1 is a parameter measured by device 1, and p2-1 to p2-2 are parameters related to device 2. The same is true for the other parameters p3-1 to p7-2. p1 and p2-1 are connected by an arrow. This indicates that p1 and p2-1 have a relationship, and the direction of the arrow indicates the direction of the relationship between the parameters. In the following, the arrow may be called a path. In the case of p1 and p2-1 in FIG. 3, p1 is the cause and p2-1 is the result. In addition, the number 0.6 is displayed near the arrow connecting p1 and p2-1, which indicates the strength of the relationship. The strength of the relationship ranges from 0 to 1, with 1 indicating the strongest relationship.For example, the strength of the relationship between p1 and p5-2, which are indirectly related, may be calculated by multiplying the strength of the relationship between p1 and p2-1, the strength of the relationship between p2-1 and p2-2, and the strength of the relationship between p2-2 and p5-2, and calculating the result as 0.6×0.x×0.x. This allows the strength of the relationship between parameters that are only indirectly related on the network diagram to be evaluated. According to the network diagram illustrated in FIG. 3, the direction and strength of the relationship between the parameters can be confirmed. For example, it can be confirmed that the parameter p5-2 of the device 5 has a direct or indirect relationship with all other parameters p1 to p7-2 except for itself, and also has a direct relationship with the parameter p7-2 of the indirectly connected upstream device 7. In general correlation analysis, the relationship between p5-2 and p7-2 cannot be evaluated unless the user intentionally specifies the parameters p5-2 and p7-2 and analyzes the relationship. When there are several thousand or more parameters, and one tries to analyze all of the parameter relationships between devices located far apart in a system, it takes a huge amount of work just to select the parameters to be analyzed, and it is easy to miss some parameters. In contrast, the analysis device 10 automatically analyzes the relationships between multiple parameters, targeting several thousand or more parameters, regardless of the distance between the devices or whether they are upstream or downstream. This makes it possible to analyze the relationships between a huge number of parameters without missing any parameters in a short amount of time.
[0013] In order to create a network diagram as shown in Fig. 3, the analysis device 10 applies a network analysis method using machine learning or the like to calculate correlation coefficients between parameters, connection strengths, and the like. For example, for each of several thousand or more parameters, network analysis is performed on time-series data sampled at one-second intervals for several hours or more (e.g., six hours). In this case, the following problems arise due to the characteristics of data from large-scale plants. (A) When the parameters to be analyzed in network analysis have a time lag or a time constant, taking a direct correlation with the time series data reduces the accuracy of the analysis. (B) Data acquired from plants with many unsteady phenomena, such as combustion plants, often contains unsteady data, and network analysis cannot be performed as is. (C) If the number of parameters used in network analysis increases, the resulting network diagram will become more complex, making it difficult to evaluate the relationships. Therefore, in this embodiment, pre-processing is performed on the time series data according to the characteristics of each data. As shown in Fig. 1, an analysis device 10 includes a data acquisition unit 11, an input reception unit 12, a pre-processing unit 13, an analysis unit 14, a drawing unit 15, and a storage unit 16.
[0014] The data acquisition unit 11 acquires data (referred to as target data) that is measured or the like with respect to a plurality of parameters for which a relationship is to be analyzed. For example, the target data includes measured values of parameters such as temperature, pressure, flow rate, speed, vibration, rotation speed, current, and voltage measured by various sensors installed in the plant, values calculated using these measured values, control signals such as start / stop of equipment installed in the plant and opening / closing of valves, image data, voice data, and text data acquired by a monitoring camera or the like installed in the plant. The data acquisition unit 11 stores the acquired target data in the storage unit 16. In the network analysis used to analyze the relationship between parameters, numerical data is used as learning data. The data acquisition unit 11 converts control signals such as start / stop into 0 / 1 data, for example, 0 (stop, valve closed) or 1 (start, valve open), so that the data can be applied to the network analysis, and digitizes the image data, voice data, and text data using a machine learning method such as clustering, and stores the data in the storage unit 16.
[0015] The input receiving unit 12 receives instruction information, setting information, and the like input by a user using an input device such as a keyboard, a mouse, or a touch panel.
[0016] The preprocessing unit 13 judges whether the target data can be analyzed for the relationship, and performs preprocessing on the data that requires preprocessing for the relationship analysis so that the relationship analysis can be performed. For example, when the target data A and the target data B are parameters measured at separate locations, and there is a time lag between when a characteristic value is measured for the upstream target data A and when a value having the same characteristic is measured for the downstream target data B, a process is performed to remove the effect of this time lag. For example, the preprocessing unit 13 performs statistical processing on a data group including the target data A and B to detect the presence of a time lag between the target data A and B. For example, a method may be used in which the average value and the variation are calculated for each predetermined time frame for the time series data of each target data, and the time changes of the calculated values are compared between the target data. When the presence of a time lag between the target data A and B is detected, the preprocessing unit 13 performs frequency analysis on the target data A and B by fast Fourier transform (FFT) or the like. When peaks are observed at the same frequency for the target data A and B, the relationship between the target data A and B can be evaluated by focusing on this frequency component regardless of the length of the time lag. The preprocessing unit 13 may analyze the time delay from when a characteristic value is measured for the target data A until the same characteristic value is measured for the target data B by using dynamic time warping (DTW) or a cross-correlation function. If the time delay can be analyzed, for example, the preprocessing unit 13 processes the data by shifting the time axis of the target data A or the target data B so that the relationship between the target data A and B can be evaluated. In the above (A) when there is a time delay in the parameters to be subjected to network analysis, the analysis accuracy decreases when a direct correlation is taken for the time series data, and the preprocessing unit 13 deals with this by converting to a frequency problem by FFT or removing the time delay based on an evaluation of the similarity of the waveform and fluctuation of the data using DTW.
[0017] The preprocessing unit 13 judges whether the target data is data of a stationary process. For example, the preprocessing unit 13 performs statistical processing on the target data (time series data) to calculate an average value for each predetermined time frame, and judges the target data to be data of a stationary process if the fluctuation of the average value falls within a predetermined range, and judges the target data to be data of a stationary process if the average value fluctuates greatly. If the target data is not data of a stationary process, the nature of the data changes depending on whether the target data is viewed in the long term or in the short term. Therefore, if the target data is not data of a stationary process, a process of converting the target data to data of a stationary process, such as a difference process, is performed. The difference process is a process of calculating the difference between a value of the target data at a certain time and a value a predetermined time ago. The preprocessing unit 13 performs the difference process over the entire region in the time direction of the target data, which is time series data. For example, the difference between the data 5 seconds after the beginning of the target data and the first data is calculated, the difference between the data 6 seconds after the beginning and the data 1 second after the beginning is calculated, and so on until the value of the last time of the time series data. When target data A is data that is not a stationary process and has a certain trend, and target data B is data that is a stationary process, the trend is removed from target data A by differential processing, and it is converted into stationary data. This makes it possible to evaluate the relationship between target data A and B without the influence of the trend.
[0018] When the data determined to be capable of relationship analysis and the target data after preprocessing are taken as the analysis target candidate data, the preprocessing unit 13 extracts the data to be analyzed from the analysis target candidate data if the number of analysis target candidate data (the number of parameters) is equal to or greater than a predetermined threshold. This is performed for the purpose of removing unnecessary data from the processing target, since if there is too much analysis target candidate data, it is possible that data that clearly has no relationship is included in the data, and it is preferable to remove such data from the viewpoint of processing load and processing speed, and removing data that clearly has no relationship does not affect the analysis of the relationship between parameters. Note that the data determined to be capable of relationship analysis refers to target data that is determined to be data in a steady process without a time delay. For example, the preprocessing unit 13 extracts only a predetermined number of target data in descending order of correlation coefficient or similarity based on the correlation coefficient or similarity between the analysis target candidate data.
[0019] The analysis unit 14 analyzes the relationship between the target data by network analysis. There are several methods of network analysis, such as (a) correlation and partial correlation analysis, (b) structural equation modeling (SEM), and (c) linear non-Gaussian directed model (DirectLinGAM). (a) correlation and partial correlation analysis evaluates the strength of the connection (edge) between parameters based on correlation coefficients and partial correlation coefficients, and creates a network by manually determining the direction (path) indicating the cause and result between parameters. It can be used when the strength of the connection between parameters does not require consideration of upstream influences. (b) structural equation modeling estimates the connection and direction between parameters in advance by a person, and creates a network by evaluating the estimated results using machine learning. Unlike correlation and partial correlation analysis, it evaluates the strength of the connection by considering upstream parameters based on the function form, and can delete connections between parameters with low correlation by evaluating the independence of the parameters. (c) linear non-Gaussian directed model can automatically evaluate the strength of the relationship between parameters by considering the influence of upstream parameters, as with structural equation modeling. Furthermore, in the linear non-Gaussian directed model, the directionality between data can be automatically calculated to create a network. By changing the network analysis method, the characteristics of the output network diagram (for example, FIG. 3) can be changed. For example, in the case of the (c) linear non-Gaussian directed model, which evaluates the strength of the relationship between parameters taking into account the influence of upstream parameters, the relationship between parameters can be organized not by correlation but by directionality, causality, and the strength of the relationship, and a highly explanatory network diagram can be created. The analysis unit 14 evaluates the relationship between parameters using (a) to (c) or other network analysis methods. In this specification, the process when the (c) linear non-Gaussian directed model is applied will be mainly described.
[0020] The drawing unit 15 creates a network diagram, a pie chart, etc. based on the analysis result of the relationship by the analysis unit 14, and outputs the created diagram or graph to a display device, an electronic file, etc. For example, the drawing unit 15 creates a network diagram (e.g., FIG. 3) that includes information indicating the presence or absence of a relationship between parameters, the strength of the relationship such as a correlation coefficient or connection strength, and information indicating the directionality of the relationship (cause, result), and also takes into account parameters on the upstream side of the plant system. By referring to the network diagram, for example, when an abnormal value is measured for a certain parameter, it becomes easier to identify other parameters that may have influenced the parameter. In addition, when a network analysis is performed, information indicating the strength of the relationship between the parameters is output. The drawing unit 15 organizes the influence of each parameter on the target data based on the information indicating the strength of the relationship, and outputs it as a pie chart as exemplified in FIG. 4. The pie chart as exemplified in FIG. 4 is a graph that focuses on the strength of the relationship between parameter p5-2 and other parameters that have a relationship with p5-2, and distributes the weight of each parameter according to the value of the relationship strength so that the total is 100%. For example, p7-2 is a parameter that has not been recognized before. In this way, according to the relationship analysis of this embodiment, the user does not select the parameters for which the relationship is to be analyzed, but the relationship is analyzed for all parameters, so that a parameter having a relationship that has not been recognized before can be discovered, and a relationship with a parameter related to a device that is located at a distant position in the system can be confirmed.
[0021] The storage unit 16 stores various pieces of information acquired by the data acquisition unit 11 and the input reception unit 12. The storage unit 16 may be configured to include an external storage device such as a so-called cloud computing system.
[0022] (operation) Next, a process for analyzing the relationships among a huge number of parameters will be described with reference to Figures 5 and 6. Figure 5 is a flowchart showing an example of the process for analyzing the relationships according to the embodiment. First, the data acquisition unit 11 acquires data on a large number of parameters collected in a large-scale plant or the like (step S1). For example, there may be several thousand types of parameters, and the acquired data may include several hours' worth of data sampled at one-second intervals (time-series data). The data acquisition unit 11 stores the acquired data in the storage unit 16 and outputs it to the pre-processing unit 13. The data acquisition unit 11 converts control signals and image data into digitized data, adds time information to convert the data into time-series data, and then stores the data in the storage unit 16 or outputs the data to the pre-processing unit 13.
[0023] Next, the preprocessing unit 13 executes preprocessing (step S2). FIG. 6 shows an example of preprocessing. The preprocessing unit 13 calculates an autocorrelation coefficient for each parameter data acquired in step S1 (step S201), and calculates the time required for learning using the calculated autocorrelation coefficient and Bartlett's formula (step S202). This calculates how many hours of target data are required as learning data required for network analysis. The preprocessing unit 13 determines whether time-series data has been acquired over a period of time equal to or longer than the period calculated in step S202 (step S203). If the time-series data is insufficient (step S203; No), the preprocessing unit 13 adds time-series data so that the period of time is equal to or longer than the period calculated in step S202 (step S204). For example, the preprocessing unit 13 adds time-series data by the same process as step S1. If sufficient time-series data has been acquired (step S203; Yes), the preprocessing unit 13 proceeds to step S205.
[0024] In step S205, the preprocessing unit 13 determines whether or not there is a time delay between the multiple target data (step S205). If there is a time delay (step S205; Yes), the process proceeds to step S210. In a large-scale plant, if there is no time delay between the target data (step S205; No), the process proceeds to the preprocessing shown in steps S206 and after. Although there are few cases where there is no time delay, the preprocessing shown in steps S206 and after can improve the accuracy of the analysis.
[0025] In step S206, the preprocessing unit 13 judges whether the data is a stationary process (step S206). If the data is a stationary process (step S206; Yes), the preprocessing unit 13 judges whether to perform frequency analysis on the entire target data (step S207). For example, the frequency analysis method may be FFT. At this time, whether to perform frequency analysis can be arbitrarily selected depending on whether the frequency analysis improves the accuracy and ease of analysis of the target data. This selection may be performed by the user each time, or may be selected based on a setting in advance as to whether or not to perform frequency analysis on the entire target data. If it is determined that the frequency analysis is to be performed on the entire target data (step S207; Yes), the process proceeds to step S210. If it is determined that the frequency analysis is not to be performed on the entire target data (step S207; No), the process proceeds to step S211. If the data is not a stationary process (step S206; No), the preprocessing unit 13 judges whether to perform frequency analysis on the entire target data (step S208). In this case, whether or not to perform frequency analysis can be arbitrarily selected depending on whether or not the accuracy and ease of analyzing the target data are improved by frequency analysis. If it is determined that frequency analysis is to be performed on the entire target data (step S208; Yes), the process proceeds to step S210. If it is determined that frequency analysis is not to be performed on the entire target data (step S208; No), the preprocessing unit 13 performs differential processing (step S209) and proceeds to step S211.
[0026] In step S210, if the result of step S205, step S207, or step S208 is Yes, the preprocessing unit 13 performs frequency analysis on the entire target data (step S210). One of the purposes of performing frequency analysis is to remove the effect of time lag. The method of removing the effect of time lag is not limited to frequency analysis such as FFT, and the effect of time lag can also be removed by analyzing the delay time using DTW or cross-correlation function, and shifting the time axis of one of the data by the amount of the delay time. In this case, it is preferable to check whether the target data follows a normal distribution (step S211). For example, when the effect of time lag is removed using DTW or cross-correlation function, it is possible to check whether the target data follows a normal distribution after step S210 (step S211), and if it does not follow the normal distribution, to perform the process of step S212, and then proceed to step S213.
[0027] In step S211, if it is determined that the result of step S207 is No, or after step S209, the preprocessing unit 13 judges whether the target data follows a normal distribution (step S211). At this time, a method for judging whether the target data follows a normal distribution includes a QQ (Quantile-Quantile) plot. If the target data follows a normal distribution (step S211; Yes), the process proceeds to step S213. If the target data does not follow a normal distribution (step S211; No), the target data is transformed so that it follows a normal distribution (step S212). Methods for transforming the data so that it follows a normal distribution include Yeo-Johnson transformation and Box-Cox transformation.
[0028] In step S213, the preprocessing unit 13 determines whether the number of data to be analyzed (the number of parameters) is 300 or less (step S213). The data to be analyzed is either target data that follows a normal distribution, target data from which the effect of time delay has been removed, or target data that has been transformed to follow a normal distribution by Yeo-Johnson transformation or the like, with respect to the data acquired in step S1. The number 300 is just an example. For example, if the relationship between parameters is to be analyzed using this embodiment when a problem occurs, not much time can be spent on it. Therefore, it is considered to narrow down the number of data to be analyzed to, for example, about 300. If the number is 300 or less (step S213; Yes), the preprocessing is terminated. If the number of parameters exceeds 300 (step S213; No), the preprocessing unit 13 calculates the correlation coefficient between the target parameters and the like (step S214). The preprocessing unit 13 calculates the correlation coefficient and the similarity of the data by machine learning or the like. Then, the preprocessing unit 13 extracts the top 300 data items that are strongly related to the parameters of the target to be analyzed by arranging the calculated correlation coefficients or similarities in descending order (step S215). For example, a user inputs the parameters of the target to be analyzed to the analysis device 10. The input receiving unit 12 acquires the input parameters. The preprocessing unit 13 extracts 300 or less parameters that are strongly related to the input parameters. This ends the preprocessing. The preprocessing performs a time lag removal process as necessary, and extracts a desired number (for example, 300 or less) of target data that is ready for network analysis, making it possible to perform analysis in a realistic time and with a realistic accuracy. The preprocessing unit 13 stores the target data obtained by the preprocessing in the storage unit 16 and outputs it to the analysis unit 14.
[0029] Returning to FIG. 5, the analysis unit 14 performs a network analysis on the target data that has passed the preprocessing (step S3). For example, the analysis unit 14 performs an analysis applying a linear non-Gaussian directed model. Next, the analysis unit 14 deletes the remaining paths from the results of the network analysis, leaving only paths that are directly or indirectly connected to the parameter to be analyzed (step S4). For example, in the example of FIG. 3, if the parameter to be analyzed is p5-2, p7-2 and p7-3 (not shown) have a relationship, the direction of the relationship is that p7-2 is the cause and p7-3 is the result, and p7-3 has no relationship with p5-2, the analysis unit 14 deletes the arrow (path) between p7-2 and p7-3.
[0030] Next, the analysis unit 14 sets a threshold value for the connection strength of the path (step S5). For example, the analysis unit 14 sets 0.1 as the initial value of the threshold. The threshold value for the connection strength may be set by the user. The user inputs a threshold value to the analysis device 10, and the input reception unit 12 acquires the input threshold value. The analysis unit 14 sets the input value as the threshold value. Next, the analysis unit 14 extracts parameters whose connection strength with the parameter to be analyzed is equal to or greater than the threshold value. The connection strength is a value such as 0.6 in FIG. 3. For parameters having an indirect relationship, the analysis unit 14 calculates the connection strength by multiplying the connection strength with the parameter that exists between them. If there are more than 50 parameters whose connection strength is equal to or greater than the threshold value (step S6; No), the analysis unit 14 increases the threshold value by 0.05 (step S7) and extracts parameters whose connection strength is equal to or greater than the threshold value (step S5). This reduces the number of parameters to be extracted. The parameters extracted here are the parameters that are ultimately output to the network diagram. If there are 50 or less parameters whose connection strength is equal to or greater than the threshold (step S6; Yes), the analysis unit 14 judges whether there are 10 or less parameters whose connection strength is equal to or greater than the threshold, and if there are 10 or less parameters (step S8; Yes), the threshold is reduced by 0.05 (step S9), and parameters whose connection strength is equal to or greater than the threshold are extracted based on the reduced threshold (step S5). This ensures that the number of extracted parameters is 10 or more. If there are 10 or more and 50 or less parameters whose connection strength is equal to or greater than the threshold, the analysis unit 14 outputs the extracted parameters to the drawing unit 15. Note that the values of 10, 50, 0.05, etc. for the number of parameters to be extracted are merely examples, and can be set arbitrarily according to the operation.
[0031] Next, the drawing unit 15 creates and outputs a network diagram (step S10). For example, when p5-2 is specified as the parameter to be analyzed, the drawing unit 15 creates a network diagram as exemplified in FIG. 3. The drawing unit 15 may create and output not only a network diagram but also a pie chart as exemplified in FIG. 4. It takes a lot of man-hours to manually create a network diagram of a large-scale plant having large-scale data. In contrast, by appropriately naming and tagging each parameter and applying a machine learning method such as morphological analysis to the character strings of the names and tags, it is possible to group each parameter according to the system in the plant, rearrange them in a desired order, and extract only parameters having a predetermined character string, and it becomes possible to automatically output a network diagram according to the results. For factor analysis and correlation evaluation by experts, it is desirable to output a network diagram that takes the system into consideration. For example, if each parameter can be grouped according to the system in the plant, it is possible to extract parameters by system and create a network diagram by system.
[0032] (effect) As described above, according to this embodiment, instead of evaluating one-to-one relationships between parameters selected by a user, network analysis is applied to large-scale data, making it possible to automatically evaluate the relationships between each parameter and other parameters in a brute-force manner for all parameters. This eliminates the need for a user to select a parameter to be analyzed, and makes it possible to automatically analyze relationships between a huge number of parameters regardless of the number of parameters to be analyzed. In addition, it is possible to automatically analyze relationships with parameters upstream in the system that are not directly related to the parameter to be analyzed.
[0033] FIG. 7 is a diagram showing an example of a hardware configuration of an analysis device according to each embodiment. The computer 900 includes a CPU 901, a main storage device 902, an auxiliary storage device 903, an input / output interface 904, and a communication interface 905. The above-mentioned analysis device 10 is implemented in the computer 900. The above-mentioned functions are stored in the auxiliary storage device 903 in the form of a program. The CPU 901 reads the program from the auxiliary storage device 903, loads it in the main storage device 902, and executes the above-mentioned processing according to the program. The CPU 901 also secures a storage area in the main storage device 902 according to the program. The CPU 901 also secures a storage area in the auxiliary storage device 903 for storing data being processed according to the program.
[0034] A program for realizing all or part of the functions of the analysis device 10 may be recorded on a computer-readable recording medium, and the program recorded on the recording medium may be read into a computer system and executed to perform processing by each functional unit. The term "computer system" as used herein includes hardware such as an OS and peripheral devices. In addition, if a WWW system is used, the term "computer system" also includes a homepage providing environment (or display environment). In addition, the term "computer-readable recording medium" refers to a portable medium such as a CD, DVD, or USB, or a storage device such as a hard disk built into a computer system. In addition, if the program is distributed to the computer 900 via a communication line, the computer 900 that receives the program may deploy the program in the main storage device 902 and execute the above processing. In addition, the above program may be for realizing part of the above-mentioned functions, or may be capable of realizing the above-mentioned functions in combination with a program already recorded in the computer system.
[0035] As described above, several embodiments according to the present disclosure have been described, but all of these embodiments are presented as examples and are not intended to limit the scope of the invention. These embodiments can be implemented in various other forms, and various omissions, substitutions, and modifications can be made without departing from the gist of the invention. These embodiments and their modifications are included in the scope of the invention and its equivalents as described in the claims, as well as in the scope and gist of the invention.
[0036] <Additional Notes> The analysis device, analysis method, and program described in each embodiment can be understood, for example, as follows.
[0037] (1) The analysis device 10 of the first aspect includes a data acquisition unit that acquires data for each of a plurality of parameters for which relationships are to be analyzed; a pre-processing unit that determines whether the data can be analyzed for relationships and, for the data that needs to be processed for the analysis of the relationships, performs pre-processing so that the relationships can be analyzed; an analysis unit that analyzes the relationships of the data that is determined to be capable of being analyzed and the data that has been pre-processed; and an output unit that outputs information indicating the relationships of the parameters based on the results of the analysis. This makes it possible to analyze the relationships between parameters based on large-scale data containing a large number of parameters, without the user having to select the parameters to be analyzed for the relationships. It is also possible to analyze the relationships between the target data and parameters of equipment, etc., located upstream or downstream of the system that do not have a direct relationship with the target data.
[0038] (2) An analysis device according to a second aspect is the analysis device of (1), wherein the data is time series data, and the pre-processing unit performs processing to remove the effects of the time delay when there is a time delay between the data to be analyzed. This makes it possible to analyze the relationship between data that has a time lag.
[0039] (3) An analysis device according to a third aspect is the analysis device of (2), wherein the pre-processing unit executes FFT (fast Fourier transform), DTW (Dynamic Time Warping), or a cross-correlation function as a process for removing the effect of the time delay. This makes it possible to analyze the relationship between data that has a time lag.
[0040] (4) An analytical device according to a fourth aspect is the analytical device of (1) to (3), wherein the data is time series data, and the pre-processing unit determines whether the data is data of a steady process, and if the data is not data of the steady process, performs differential processing to calculate the difference between a value of the data at a first time and a value a predetermined time before the first time, from the value at the first time to the value at the last time of the data. This makes it possible to correct steady-state data into data that allows for relationship analysis.
[0041] (5) The analytical apparatus according to a fifth aspect is the analytical apparatus of (4), wherein the preprocessing unit determines whether the data follows a normal distribution when the data is in a stationary process or when the differential processing has been performed, and if the data does not follow the normal distribution, converts the data to follow a normal distribution by performing a Yeo-Johnson transformation, a Box-Cox transformation, or the like. This makes it possible to improve the accuracy of analysis of relationships between data.
[0042] (6) An analysis apparatus according to a sixth aspect is the analysis apparatus of any one of (1) to (5), further comprising an extraction unit that, when the data determined by the pre-processing unit to be capable of analysis and the data that has been pre-processed are taken as candidate data to be analyzed, extracts data to be analyzed from the candidate data to be analyzed if the number of the candidate data to be analyzed is equal to or greater than a predetermined threshold, and the extraction unit extracts the data to be analyzed based on a correlation coefficient or a similarity between the candidate data to be analyzed. This allows data that is clearly not relevant to be removed.
[0043] (7) An analysis device according to a seventh aspect is an analysis device according to any one of (1) to (6), wherein the analysis unit analyzes the relationships between the data by network analysis, and extracts the data whose connection strength indicating the strength of the relationship between the data included in the results of the network analysis is equal to or greater than a predetermined threshold value. This makes it possible to extract only data with strong correlations.
[0044] (8) An analysis device according to an eighth aspect is the analysis device of (7), wherein when the number of the data whose connection strength is equal to or greater than the threshold exceeds a predetermined first number, the analysis unit increases the threshold of the connection strength by a predetermined value and extracts the data based on the increased threshold. This makes it possible to extract data in order of strongest correlation.
[0045] (9) An analysis device according to a ninth aspect is the analysis device of (7) to (8), wherein when the number of the data whose connection strength is equal to or greater than the threshold is equal to or less than a predetermined second number, the analysis unit reduces the threshold for the connection strength by a predetermined value and extracts the data based on the reduced threshold. In this way, if there is not enough data with a strong relationship, it is possible to extract data with the next strongest relationship.
[0046] (10) An analysis device according to a tenth aspect is an analysis device according to any one of (1) to (9), wherein the output unit outputs a network diagram including the strength and direction of the relationship between the parameters based on the analysis results of the relationship between the data. This makes it possible to visualize the relationships between parameters.
[0047] (11) An analysis device according to an eleventh aspect is an analysis device according to any one of (1) to (10), wherein the output unit outputs a graph showing the degree of strength of the relationship between one of the parameters and each of the other parameters having a relationship with the one of the parameters based on the strength of the relationship between the one of the parameters. This makes it possible to visualize, when focusing on one parameter, the strength of other parameters that have a relationship with that parameter.
[0048] (12) An analysis method according to a twelfth aspect is an analysis method executed by a computer, which acquires data for each of a plurality of parameters for which relationships are to be analyzed, determines whether or not analysis of the relationships is possible for the data, performs preprocessing on the data that requires processing for the relationship analysis so that the relationship can be analyzed, analyzes the relationships of the data that is determined to be capable of being analyzed and the data that has been preprocessed, and outputs information indicating the relationships of the parameters based on the results of the analysis.
[0049] (13) A program according to a thirteenth aspect causes a computer to function as: means for acquiring data for each of a plurality of parameters for which relationships are to be analyzed; means for determining whether or not relationship analysis is possible for the data, and for the data that requires processing for the relationship analysis, performing preprocessing so as to enable relationship analysis; means for analyzing the relationships between the data for which it is determined that the analysis is possible and the data that has been preprocessed; and means for outputting information indicating the relationships between the parameters based on the results of the analysis. [Explanation of symbols]
[0050] 10...Analyzer 11 Data acquisition section 12 Input reception section 13. Pretreatment section 14...Analysis Department 15. Drawing section 16...Storage section 1,2,3,4,5,6,7...Equipment 100···Plant 900...Computer 901··CPU 902...Main memory 903...Auxiliary storage device 904 Input / Output Interface 905 Communication Interface
Claims
1. A data acquisition unit that acquires data of each of a plurality of parameters for analyzing a relationship; a pre-processing unit that determines whether or not a relationship analysis of the data is possible, and performs pre-processing on the data that needs to be processed for the relationship analysis so that the relationship analysis is possible; an analysis unit that analyzes the relationship between the data that is determined to be analyzable and the data that has been preprocessed; an output unit that outputs information indicating the relationship between the parameters based on a result of the analysis; An analytical device comprising:
2. The data is time series data, The preprocessing unit performs a process for removing an effect of the time delay when there is a time delay between the data for which the relationship is to be analyzed. The analytical device of claim 1 .
3. The pre-processing unit executes FFT (fast Fourier transform), DTW (Dynamic Time Warping), or a cross-correlation function as a process for removing the effect of the time delay. The analytical device according to claim 2 .
4. The data is time series data, the preprocessing unit determines whether the data is data of a steady process, and if the data is not data of the steady process, performs a difference process for calculating a difference between a value of the data at a first time and a value a predetermined time before the first time, from a value at a first time to a value at a last time of the data; The analysis device according to claim 1 or 2.
5. When the data is a stationary process data or when the difference processing is performed, the preprocessing unit determines whether the data follows a normal distribution, and performs a Yeo-Johnson transformation or a Box-Cox transformation on the data that does not follow the normal distribution. The analytical device according to claim 4.
6. an extraction unit that extracts data to be analyzed from the analysis target candidate data when the number of the analysis target candidate data is equal to or greater than a predetermined threshold value when the data determined by the preprocessing unit to be capable of analysis and the data that has been preprocessed are regarded as analysis target candidate data; Further equipped with The extraction unit extracts the data to be analyzed based on a correlation coefficient or a similarity between the analysis target candidate data. The analysis device according to claim 1 or 2.
7. the analysis unit analyzes the relationship between the data by network analysis, and extracts the data having a connection strength indicating a strength of the relationship between the data included in a result of the network analysis that is equal to or greater than a predetermined threshold value; The analysis device according to claim 1 or 2.
8. When the number of the data whose connection strength is equal to or greater than the threshold exceeds a predetermined first number, the analysis unit increases the threshold of the connection strength by a predetermined value and extracts the data based on the increased threshold. The analytical device according to claim 7.
9. When the number of the data whose connection strength is equal to or greater than the threshold is equal to or less than a predetermined second number, the analysis unit reduces the threshold of the connection strength by a predetermined value and extracts the data based on the reduced threshold. The analytical device according to claim 7.
10. The output unit outputs a network diagram including a strength of a relationship and a directionality of the relationship between the parameters related to the data based on an analysis result of the relationship between the data. The analysis device according to claim 1 or 2.
11. the output unit outputs a graph indicating a degree of strength of a relationship between one of the parameters and each of the other parameters having a relationship with the one of the parameters, based on the strength of the relationship between the one of the parameters. The analytical device of claim 10.
12. 1. A computer-implemented analysis method comprising: Obtain data for each of the multiple parameters whose relationships are to be analyzed, determining whether a relationship analysis is possible for the data, and performing preprocessing on the data that needs to be processed for the relationship analysis so that the relationship analysis is possible; Analyzing the relationship between the data that is determined to be analyzable and the data that has been pre-processed; outputting information indicating the relationship between the parameters based on the results of the analysis; Analysis method.
13. Computer, A means for acquiring data for each of the multiple parameters for which relationships are to be analyzed; means for determining whether or not a relationship analysis is possible for the data, and for the data that needs to be processed for the relationship analysis, performing pre-processing so that the relationship analysis is possible; A means for analyzing the relationship between the data determined to be analyzable and the data that has been pre-processed; means for outputting information indicating the relationship between the parameters based on the results of the analysis; A program to function as a
Citation Information
Patent Citations
Abnormality determination device, learning device, and abnormality determination method
JP2022074890A