DATA SELECTION PROCESSING SYSTEM AND DATA SELECTION PROCESSING METHOD

The data selection processing system addresses the inefficiencies in existing sensor group definition methods by using correlation coefficients to dynamically adjust sensor group sizes and numbers, enhancing diagnostic accuracy and reducing processing loads in large-scale systems.

JP7672952B2Active Publication Date: 2025-05-08HITACHI LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2021182545
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-11-09
Publication Date
2025-05-08
Estimated Expiration
2041-11-09

AI Technical Summary

Technical Problem

Existing methods for defining sensor groups in large-scale systems like plants are inefficient, as they often result in biased sensor group sizes and excessive processing loads, requiring specialized knowledge and not allowing for adjustment of sensor numbers in each group.

Method used

A data selection processing system that uses a sensor data acquisition unit, an analysis unit to calculate correlation coefficients, and an analysis group creation unit to define sensor groups based on high correlation coefficients, allowing for adjustment of sensor numbers in each group and the total number of sensors used.

Benefits of technology

Enables the formation of appropriate sensor groups, allowing for improved diagnostic accuracy by adjusting sensor group sizes and reducing processing loads, without requiring specialized knowledge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007672952000001
    Figure 0007672952000001
  • Figure 0007672952000002
    Figure 0007672952000002
  • Figure 0007672952000003
    Figure 0007672952000003
Patent Text Reader

Abstract

To configure groups with an appropriate number of sensors when grouping sensors for analysis.SOLUTION: A data selection processing system includes; a sensor data acquisition unit that acquires detection data from a plurality of sensors; an analysis unit that obtains correlation coefficients of detection data from the plurality of sensors captured by the sensor data acquisition unit; and an analysis group creation unit that defines as a group a set of sensors whose correlation coefficient obtained by the analysis unit is higher than a threshold value mutually set between sensors.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a data selection processing system and a data selection processing method. [Background technology]

[0002] Large-scale systems such as plants are equipped with many sensors, numbering up to several thousand. When performing anomaly diagnosis processing on such large-scale systems, instead of performing a single analysis process that inputs all the sensor data, a method is often adopted in which groups consisting of a small number of sensors divided into units of systems or equipment are defined, and analysis processing for each group is performed simultaneously in parallel. This is because by dividing the diagnosis processing of a large-scale system into multiple groups and analyzing them, it is possible to identify the system or equipment in which an anomaly has occurred as a group. Furthermore, by analyzing in units of groups consisting of a small number of sensors, the sensitivity to the analysis results due to a change in the value of a single sensor is increased, and the accuracy of anomaly detection can be improved.

[0003] However, the task of defining groups suitable for diagnosis consisting of a small number of sensors in units of systems or equipment for a large amount of sensor data must be determined in consideration of the equipment configuration of the target of diagnosis, such as the plant, and the impact on each sensor of various abnormalities. Therefore, the definition of groups suitable for diagnosis requires specialized knowledge, which has been a bottleneck when starting diagnostic processing. Given this background, there was a demand for a processing method that could automatically define groups from sensor data without using specialized knowledge of the plant being diagnosed.

[0004] In response to such demands, one possible method is to define a group of sensors that have a high correlation (correlation coefficient) between them. This method relies on the fact that sensors installed in the same system or equipment have a high correlation coefficient. For example, in the case of a flow sensor, if sensors are installed on the upstream and downstream sides of the same fluid path, the measured values ​​of the respective flow rates change in the same way, and therefore show a high correlation. In addition, if a pump is taken as an example of a sensor installed in the same equipment, if a control method is adopted that adjusts the pump discharge flow rate according to the rotation speed of the drive motor, the measured values ​​of the rotation speed and the discharge flow rate show a high correlation.

[0005] As described above, defining a group by collecting sensors that are highly correlated with each other based on the correlation coefficient between sensors is an effective method for automatically determining groups without using specialized knowledge.

[0006] Patent Document 1 describes a technology for analyzing data composed of multiple sensor signals. In the example described in Patent Document 1, a directed acyclic graph, which is a type of directed graph, is created from the correlation coefficient between sensors, and the root cause of an anomaly is identified using the relationship between the nodes shown in the graph. A directed graph is one in which the edges connecting the nodes are provided with directional information (vectors). This makes it possible to graphically express causal relationships, such as the effect of a specific sensor on another sensor. [Prior art documents] [Patent documents]

[0007] [Patent Document 1] Patent Publication No. 2021-2335 Summary of the Invention [Problem to be solved by the invention]

[0008] However, known methods such as the technique described in Patent Document 1 have a problem in that the number of sensors in each group cannot be adjusted. Thousands of sensors are often installed in large-scale systems such as plants. When groups are defined using known methods for data consisting of such a large number of sensors, in extreme cases, the number of sensors may be significantly biased depending on the group, such as one group having thousands of sensors and another group having only a few sensors. This makes it difficult to achieve the original purpose of defining multiple analysis groups to improve diagnostic accuracy.

[0009] In addition, there are cases where it is not necessary to use all of the thousands of sensors installed in a large-scale system in diagnostic processing. This is because if all sensors are used, many groups will be defined, but processing of all groups may not be possible due to the processing load of the system. Furthermore, it may become difficult to know which group's analysis results to pay attention to, making it more difficult to understand the plant status.

[0010] As described above, there has been a demand for a method of using the correlation coefficient of each sensor to represent the correlation as a graph, and when highly correlated sensors are grouped together based on the graph to define a group for analysis, which allows the number of sensors in each group and the number of sensors used for analysis to be adjustable, or a method of allowing the number of sensors in each group and the number of groups created to be adjustable.

[0011] An object of the present invention is to provide a data selection processing system and a data selection processing method that, when grouping sensors to be used for analysis, can form groups with an appropriate number of sensors. [Means for solving the problem]

[0012] In order to solve the above problems, for example, the configurations described in the claims are adopted. The present application includes multiple means for solving the above problems. One example is a data selection processing system that includes a sensor data acquisition unit that acquires detection data from multiple sensors, an analysis unit that calculates a correlation coefficient of the detection data from the multiple sensors acquired by the sensor data acquisition unit, and an analysis group creation unit that defines a group as a set of sensors for which the correlation coefficient calculated by the analysis unit is higher than or equal to a threshold value set between the sensors. Here, the analysis unit sets each sensor as a node, creates a graph in which edges are set between sensors whose correlation coefficients are above a threshold, and determines a set of sensors that are highly correlated with each other based on a clique or clique community. Effect of the Invention

[0013] According to the present invention, when defining a group that groups together sensors that are highly correlated, the number of sensors in each group and the total number of sensors used in analysis can be adjusted to define the groups. Problems, configurations and effects other than those described above will become apparent from the following description of the embodiments. [Brief description of the drawings]

[0014] [Figure 1] 1 is a configuration diagram showing an example of a plant diagnosis system according to a first embodiment of the present invention. [Diagram 2] FIG. 3 is a diagram showing an example of a data structure of a sensor value database according to the first embodiment of the present invention. [Diagram 3] FIG. 4 illustrates correlation coefficients in graph form for four sensors according to a first exemplary embodiment of the present invention. [Figure 4] FIG. 4 is a diagram illustrating a schematic diagram of a process for adjusting the number of sensors in each group according to the first embodiment of the present invention. [Diagram 5] 5 is a flowchart showing a process flow for defining a group according to the first embodiment of the present invention. [Figure 6] FIG. 2 is a diagram showing an example of a data structure of an analysis group database according to the first embodiment of the present invention. [Figure 7] FIG. 2 is a diagram showing an example of a data structure of an analysis result database according to the first embodiment of the present invention. [Figure 8]FIG. 11 is a diagram showing an example of a screen when creating an analysis group according to the first embodiment of the present invention. [Figure 9] FIG. 2 is a diagram showing an example of a data structure of a tag number database according to the first embodiment of the present invention. [Figure 10] FIG. 4 is a diagram showing an example of a display screen showing an analysis result of a diagnosis according to the first embodiment of the present invention. [Figure 11] 10 is a flowchart showing a process flow for defining a group according to a second embodiment of the present invention. [Figure 12] FIG. 11 is a diagram showing an example of a display screen of the system showing analysis groups according to the second embodiment of the present invention. [Figure 13] FIG. 2 is a block diagram showing an example of a hardware configuration when a system according to each embodiment of the present invention is configured using a computer. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0015] <First embodiment> A data selection processing system according to a first embodiment of the present invention will now be described with reference to FIGS.

[0016] [Configuration of a plant diagnostic system equipped with a data selection processing system] FIG. 1 shows the configuration of a plant diagnostic system equipped with a data selection processing system according to a first embodiment of the present invention. The plant diagnostic system 1 is supplied with sensor data, which is detection data from sensors installed in a plant 3. In addition, an input / output device 2 is connected to the plant diagnostic system 1. The input / output device 2 sets analysis conditions related to data selection processing for the plant diagnostic system 1, and presents the analysis results of the diagnosis. The plant 3 is a target for diagnosis by the plant diagnostic system 1, that is, a plant to which data selection processing for diagnosis is applied.

[0017] The sensor data acquisition unit 11 of the plant diagnosis system 1 performs a sensor data acquisition process for acquiring sensor data from the plant 3, and stores the acquired sensor data in the sensor value database 12. Fig. 2 shows the data structure of the sensor value database 12. In the sensor value database 12, sensor values ​​are stored in chronological order for each sensor together with date and time data. Each sensor is identified by a tag number. In the example of Fig. 2, A001, A002,... are tag numbers.

[0018] Returning to the explanation of the configuration in Fig. 1, the analysis group creation unit 13 imports data from the sensor value database 12 and creates an analysis group. The process of creating an analysis group will be described in detail later. The data of the analysis group created by the analysis group creation unit 13 is stored in the analysis group database 14. For the group to be analyzed, the analysis unit 15 imports multiple tag numbers belonging to the group from the analysis group database 14, and then imports time-series data of the sensor values ​​corresponding to each tag number from the sensor value database 12.

[0019] The analysis unit 15 uses the captured sensor value data to perform analysis processing for plant diagnosis. Here, the analysis method is not particularly limited as long as it processes a plurality of time series data as input. For example, there is a method that uses clustering, which is one of pattern recognition processing. This is based on a cluster created from sensor data under normal conditions. In other words, it is a learning type method that learns in advance with sensor data under normal conditions and uses this as a reference for diagnosis. The diagnosis processing is performed based on the distance between the sensor data and the cluster under normal conditions. In other words, the greater the distance between the sensor data and the cluster under normal conditions, the more different the data pattern under normal conditions is. Therefore, when the distance between the sensor data and the cluster in the normal state is large, the distance can be used as the degree of anomaly. For example, the analysis unit 15 takes in the sensor data, and outputs the degree of anomaly based on the distance between the sensor data and the cluster in the normal state to the analysis result database 16.

[0020] The conditions for creating a group and the conditions for analyzing a diagnosis are set by the input / output control unit 18 based on the user's operation on the input / output device 2. The data of the sensor name is stored in the tag number database 17.

[0021] [Analysis group creation process] Next, the analysis group creation process performed by the analysis group creation unit 13 will be described. In this embodiment, graph theory is applied to the analysis group creation process. Graph theory refers to a general mathematical theory for a data structure called a graph that uses nodes (vertices) and edges (branches). Graph theory can be applied by creating a graph in which nodes are represented as sensors and edges are represented as pairs of highly correlated sensors. In this case, the edges are subjected to a threshold judgment on the correlation coefficient of two sensors, and an analysis group is created if it is above the threshold.

[0022] One of the signal processing concepts in graph theory is the idea of ​​a clique. A clique is a subset of a graph in which all nodes are connected to each other by edges (this is called a complete graph). When sensor correlation is expressed using a graph, a clique represents a set of sensors that are highly correlated with each other, and can be used as a group for diagnostic processing. A clique consisting of K nodes is called a K-clique. Depending on the number of nodes, it is also called a 3-clique, 4-clique, etc.

[0023] There is a concept of a clique community as a set related to cliques. When there are two K-cliques consisting of K nodes, if they share K-1 nodes, the two cliques are considered as one set. This set is called a K-clique community. In other words, a clique community is a set of cliques, and is a set that is larger than a clique.

[0024] The analysis group creation process performed by the analysis group creation unit 13 by applying the graph theory described above will be explained below. First, time series data of each sensor is taken in from the sensor value database 12, and correlation coefficients between each sensor are calculated. The number of correlation coefficients to be calculated is the same as the number of combinations of two sensors. For example, the number of combinations of two sensors for 1000 sensors is 499,500. The number of correlation coefficients to be calculated is also the same.

[0025] Figure 3(a) shows the correlation coefficients for four sensors [1]-[4] in a simple graph format. Here, a graph refers to a data representation format that uses nodes and edges. In the example of Figure 3, sensors [1]-[4], indicated by circled numbers, are called nodes. Also, a line connecting two nodes (sensors) is called an edge. Since the number of combinations for the four sensors is six, six edges are shown.

[0026] In FIG. 3(a), the correlation coefficient corresponding to the nodes (sensors) at both ends is written for each edge. A normal correlation coefficient is a value within the range of -1 to +1. A correlation coefficient of 1 indicates that the relative changes between the two are the same. However, even if the correlation coefficient is 1, the absolute value is not necessarily the same. Also, a correlation coefficient of -1 indicates that the relative changes between the two are completely opposite. On the other hand, a correlation coefficient of 0 indicates that there is no correlation at all between the relative changes between the two. In this embodiment, positive correlation and negative correlation are treated in the same way, so the absolute value of the correlation coefficient is used.

[0027] In the example shown in FIG. 3(a), the correlation coefficient between sensors [1] and [2] is 0.9, indicating a very high correlation. On the other hand, the correlation coefficient between sensors [2] and [4] is 0.1, indicating a low correlation. In this embodiment, a threshold judgment is performed on the correlation coefficient to define a group of highly correlated sensors. Sensors below the threshold are judged to have a low correlation, and a graph is created in which the edges between these sensors are deleted.

[0028] If the threshold value is set to 0.2 in the graph shown in Figure 3(a), the graph shown in Figure 3(b) is obtained. As this graph shows, the nodes corresponding to sensors [1], [2], and [3] are connected to each other by edges, and it can be determined that this is a set of sensors with high correlation. A set in which all nodes are connected by edges in this way is called a clique. In the example shown in Figure 3(b), the clique is composed of three nodes, so it is called a 3-clique. In other words, if a graph is created based on the correlation coefficients of each sensor and a clique is extracted from the graph, it is possible to identify a set of sensors with high mutual correlation.

[0029] The above is the basis of clique analysis. In the system according to this embodiment, a set of highly correlated sensors identified based on the clique analysis is defined as one group, and an analysis for plant diagnosis is performed for each group. Furthermore, the system according to this embodiment performs a process of adjusting the number of sensors in each group for a large-scale system consisting of a large number of sensors, such as a plant.

[0030] [Adjustment of the number of sensors] Next, the process of adjusting the number of sensors, which is performed by the analysis group creation unit 13, will be described. FIG. 4 is a diagram showing a schematic diagram of a process for adjusting the number of sensors in each group. In the example of Figure 4, there are 10 sensors [1] to

[10] , and the process is shown where the number of sensors in each group is set to 3 to 4. In Figure 4, each sensor [1] to

[10] is indicated by a number in a circle.

[0031] Figure 4(a) shows the correlation of 10 sensors [1]-

[10] in a graph. In this example, sensors [5], [6], and [9] form a 3-clique. In other words, the three sensors [5], [6], and [9] are a set that are highly correlated with each other, and these three sensors are determined as a group. On the other hand, sensors [1] to [7] form a 7-clique. This deviates from the condition that the number of sensors in each group is 3 to 4. Therefore, as shown in Figure 4(b), a process is performed to extract the 7-clique and make it into a small number of groups.

[0032] Figure 4(c) shows the result of recreating the graph by adjusting the correlation coefficient threshold in an increasing direction for the sensors that make up the seven cliques in Figure 4(b). By adjusting the correlation coefficient threshold in an increasing direction, edges between sensors with low correlation are removed. In the example in Figure 4(c), sensors [1] to [3] form three cliques. This satisfies the condition for the number of sensors, so it is determined to be a group.

[0033] On the other hand, in Figure 4(c), sensors [3]-[7] form five cliques. This deviates from the condition that specifies the number of sensors in each group as 3-4. For this reason, the five cliques shown in Figure 4(d) are extracted. Next, Figure 4(e) shows the result of recreating the graph by adjusting the correlation coefficient threshold in an increasing direction for the sensors in Figure 4(d) in the same manner as in the process described above. As a result, sensors [4]-[7] form four cliques. This satisfies the condition for the number of sensors, so it is determined to be a group.

[0034] Through the above process, three groups that satisfy the condition of having 3 to 4 sensors can be defined. In this way, if there is a clique with a node count exceeding the specified value, the correlation coefficient threshold is adjusted in an increasing direction and the graph is recreated. At this time, the graph changes in a direction that deletes edges, so a clique with a node count within the specified range is generated. If no clique is generated, the threshold is adjusted again in an increasing direction. By repeating this clique generation process, the number of sensors in all defined groups can be brought within the specified range.

[0035] [Flow of analysis group creation process and sensor number adjustment process] Fig. 5 is a flow chart showing the above-mentioned process. Steps S1 to S8 in the process flow shown in Fig. 5 are processes for adjusting the number of sensors in each group, which has been described above with reference to Fig. 4. First, the analysis group creation unit 13 sets an initial value for the threshold value of the correlation coefficient (step S1). Next, the analysis group creation unit 13 creates a graph that connects edges between sensors (nodes) whose correlation coefficients are equal to or greater than the threshold value (step S2). Furthermore, the analysis group creation unit 13 performs clique analysis on the created graph (step S3). Then, the analysis group creation unit 13 extracts cliques whose number of sensors (number of nodes) is within a specified range (step S4). The analysis group creation unit 13 determines, as a group, the sensors (nodes) that belong to the cliques extracted in step S4.

[0036] Next, the analysis group creation unit 13 determines whether or not there is a clique whose number of sensors (number of nodes) exceeds a specified value (step S6). If there is a clique whose number of sensors (number of nodes) exceeds the specified value in step S6 (YES in step S6), the analysis group creation unit 13 extracts the clique whose number of sensors (number of nodes) exceeds the specified value (step S7). Then, the analysis group creation unit 13 increases the threshold value of the correlation coefficient (step S8). Next, returning to step S2, a graph is created by connecting edges between sensors (between nodes) whose correlation coefficients are equal to or greater than the threshold value. By repeating the processes in steps S2 to S8, the number of sensors in all of the defined groups will fall within the specified range.

[0037] The above is the process that has been explained using Fig. 4. Next, the process of steps S9 to S11 will be explained. Steps S2 to S8 are processes for keeping the number of sensors in each group within a specified range, while steps S9 to S11 are processes for keeping the number of sensors in use below a specified value.

[0038] If it is determined in the above-mentioned step S6 that there is no clique whose number of sensors (number of nodes) exceeds the specified value (NO in step S6), it indicates that the number of sensors in all the confirmed groups is within the specified range. In response to this, the analysis group creation unit 13 next determines whether the number of selected sensors is equal to or less than the specified value (step S9). For example, in the group confirmed in the above-mentioned process of FIG. 4, among the ten sensors, sensors [1] to

[10] , eight sensors other than sensors [8] and

[10] are used. In contrast, it is assumed that a rule is set that a group is defined with five or less sensors. Then, since the number of sensors selected in step S9 is no longer equal to or less than the specified value (NO in step S9), the analysis group creation unit 13 resets all the confirmed groups (step S10). Note that if the number of sensors selected in step S9 is equal to or less than the specified value (YES in step S9), the process ends.

[0039] Next, the analysis group creation unit 13 increases the initial value of the correlation coefficient threshold (step S11), returns to step S1, and starts the process over from the beginning. At this time, by increasing the initial value of the correlation coefficient threshold, the number of edges of the created graph is reduced from the previous time. This also makes it possible to reduce the number of sensors ultimately used in the group.

[0040] In other words, this processing flow is composed of a double loop, with the inner loop adjusting the number of sensors in each group, and the outer loop adjusting the number of sensors to be used. The groups defined by the above processing satisfy two conditions: the number of sensors in each group is within a specified range, and the number of sensors used is equal to or less than a specified value.

[0041] In the processing shown in Figures 4 and 5, the group is defined by applying the concept of a clique, but the same method can be used with a clique community. By replacing a clique with a clique community, the same purpose can be achieved. The above is the processing of the analysis group creation unit 13 shown in Fig. 1. The processing result is output to the analysis group database 14.

[0042] [Example of analysis group database configuration] Fig. 6 shows an example of the data structure of the analysis group database 14. The analysis group database 14 is composed of two types of data shown in Fig. 6(a) and Fig. 6(b). That is, as shown in Fig. 6(a), sensors belonging to each analysis group are stored in the form of tag numbers. On the other hand, as shown in Fig. 6(b), the correlation coefficient of each analysis group when each analysis group is determined is stored.

[0043] [Example of processing performed on analysis groups] When an analysis group is created by the analysis group creation unit 13, the analysis unit 15 retrieves multiple tag numbers belonging to the group to be analyzed from the analysis group database 14, and then retrieves time series data of the sensor values ​​corresponding to each tag number from the sensor value database 12.

[0044] The analysis unit 15 uses the acquired sensor value data to perform analysis for plant diagnosis. Here, the analysis method by the analysis unit 15 is not particularly limited as long as it processes a plurality of time series data as input. For example, there is a method using clustering, which is one of pattern recognition processing. This is based on a cluster created from sensor data in normal times. In other words, it is a learning type method in which learning is performed in advance using sensor data in normal times and diagnosis is performed based on this. The diagnosis processing is performed based on the distance between the sensor data and the cluster in normal times. The larger the distance, the more different the data pattern is from normal times, so a distance equal to or greater than a threshold value may be regarded as an abnormality degree. In this embodiment, the analysis unit 15 also acquires sensor data and outputs an abnormality degree to the analysis result database 16.

[0045] 7 shows an example of the data structure of the analysis result database 16. In the analysis result database 16, the analysis results (degree of anomaly) output by the analysis unit 15 are stored in chronological order in association with the date and time data of the sensor value database 12. As described above, the analysis unit 15 performs analysis for each group, and therefore the data on the degree of anomaly is also stored for each group.

[0046] [Example of a series of steps and display screens for obtaining diagnostic analysis results based on user settings] Next, a series of steps will be described in the system according to this embodiment, in which a user sets conditions for group creation and conditions for diagnostic analysis through the input / output device 2 in Fig. 1, and obtains diagnostic analysis results accordingly. Note that the input / output data is controlled by the input / output control unit 18.

[0047] FIG. 8 is an example of a screen displayed by the input / output device 2. The screen shown in FIG. 8 is an example displayed by selecting a tab 101 that displays a screen related to analysis group creation. A window for setting conditions for analysis group creation is arranged at the top of the tab 101. Specifically, the tab 101 is arranged with a window 102 for specifying a date and time range for sensor data to be used in group creation processing. The tab 101 is also arranged with a window 103 for selecting an analysis method for group creation. This window 103 allows the selection of either clique or clique community. Furthermore, the tab 101 is arranged with a window 104 for specifying the range of the number of sensors in each group described above and the upper limit of the number of sensors to be used as conditions for group creation.

[0048] When conditions are specified in each of the windows 102 to 104 and a button 105 labeled "Create Group" is pressed, the group creation process described above is executed and the processing results are output to the analysis result database 16. Next, the processing results are displayed in a result display field 106 on the screen. The signal names of each group are displayed in a list format in the result display field 106. At this time, the input / output control unit 18 converts the tag number into a sensor name. The sensor name data is stored in the tag number database 17 in FIG. 1.

[0049] Furthermore, the number of sensors in each group is displayed in the sensor number display field 107. Furthermore, the number of sensors used is displayed in the sensor number used display field 108. This number of sensors is not the total number of sensors in each group shown in the sensor number display field 107. This is because the same signal may belong to multiple groups. By displaying the number of sensors in each group in the sensor number display field 107 and the number of sensors used in the sensor number used display field 108, it is possible to confirm that the conditions set in the window 104 are satisfied. The correlation coefficient between sensors display field 109 displays the correlation coefficient when the group is confirmed. The correlation coefficient displayed on the screen means that the correlation coefficient between sensors belonging to the same group is equal to or greater than this value.

[0050] Fig. 9 shows an example of the data structure of tag number database 17. As shown in Fig. 9, tag numbers and sensor names are stored in a corresponding manner. In analysis group database 14 shown in Fig. 1, sensors belonging to each group are stored in the form of tag numbers. Input / output control unit 18 takes out this data, converts the tag numbers to sensor names based on tag number database 17, and displays them on the screen.

[0051] The system according to this embodiment is assumed to take in data online from the plant 3 and sequentially display the results of diagnostic analysis in real time. Therefore, the operations related to the creation of the analysis group described above are performed in advance before the system executes diagnostic processing.

[0052] Fig. 10 is a display screen showing the analysis results of the diagnosis. The screen shown in Fig. 10 is displayed by selecting a tab 201 for displaying a screen related to the analysis results. A list 202 for specifying an analysis group for displaying analysis results is placed at the top of the tab 201. By selecting an analysis group in the list 202, the names of the sensors in the corresponding group are displayed in an input sensor name display field 203, and the analysis results of the analysis group are displayed in an analysis result display field 204. The example of the analysis result display field 204 shown in Fig. 10 is a trend of anomalies, which is stored in the analysis result database 16 described above.

[0053] As described above, according to this embodiment, when diagnosing a large-scale system such as a plant, a set of sensors with high correlation between them can be automatically defined as a group for diagnosis processing. Furthermore, the number of sensors in each defined group can be set within a specified range, and the number of selected sensors can be set to a specified value or less. The function of adjusting the number of sensors is particularly effective when dealing with a large-scale system in which many sensors are installed.

[0054] <Second embodiment> Next, a data selection processing system according to a second embodiment of the present invention will be described with reference to FIGS. In this embodiment, similarly to the first embodiment, diagnosis is performed on a large-scale system such as a plant, and sensor selection processing is performed in the system configuration shown in FIG. The difference between the first embodiment and the second embodiment is that, instead of the process of setting the number of selected sensors to a specified value or less in the first embodiment, the process of setting the number of groups to be created to a specified range in the second embodiment. Note that the process of setting the number of sensors in each group to a specified range, which was implemented in the first embodiment, is also implemented in the second embodiment. In addition, in the flowchart of Figure 11, which explains the analysis group creation process and the sensor number adjustment process in this embodiment, the steps in which the same processes as those in the flowchart of Figure 5 are performed are given the same step numbers and duplicate explanations are omitted.

[0055] [Flow of analysis group creation process and group number adjustment process] FIG. 11 is a flowchart showing the flow of the analysis group creation process and the group number adjustment process in this embodiment. In the flowchart of Fig. 11, steps S1 to S8 are the same as those shown in the flowchart of Fig. 5. Then, if there is no clique in which the number of sensors exceeds the specified value in step S6 (NO in step S6), a process is performed to determine whether the number of created groups is within a specified range (step S9B). In step S9B, if the number of groups is within the range of the specified value (YES in step S9B), the analysis group creation process and the group number adjustment process are terminated. On the other hand, if it is determined in step S9B that the number of groups is not within the range of the specified value (NO in step S9B), the process proceeds to step S10, where the analysis groups are cleared and returned to the initial state.

[0056] Furthermore, after the process of clearing the analysis groups and returning them to the initial state is performed in step S10, the initial value of the correlation coefficient threshold is changed (step S11B). Here, if the number of groups created is below the range of the specified value and it is desired to increase the number of groups, the threshold is adjusted downward. This increases the number of cases in which the correlation between two sensors is determined to be high, and the number of edges in the graph created in step S2 increases. As a result, the number of cliques increases, and the number of groups also increases. On the other hand, if the number of groups created is above the range of the specified value and it is desired to decrease the number of groups, the threshold can be adjusted upward in the opposite manner to the above example. After the initial value of the correlation coefficient threshold is changed in step S11B, the process returns to step S1.

[0057] By carrying out the above process, the number of groups to be created can be set within a specified range. In the flowchart of Fig. 11, the number of sensors in each group is adjusted in steps S1 to S8, and the number of groups to be created is adjusted by adding the processes of steps S9B, S10, and S11B.

[0058] [Example of display screen] Fig. 12 is an example of a display screen of the system in this embodiment. In the display screen of Fig. 12, the same items as those in the display screen shown in Fig. 8 are given the same reference numerals. The screen shown in FIG. 12 has a window 104B for setting the upper and lower limits for the number of groups, instead of the window 104 in FIG. 8 for specifying the upper limit for the number of sensors to be used. The rest of the screen in FIG. 12 is the same as that in FIG. As described above, according to the second embodiment, the number of sensors in each defined group can be kept within a specified range, and the number of groups created can be kept within a specified range.

[0059] [Variations] It should be noted that the embodiment examples described so far have been described in detail in order to explain the present invention in an easily understandable manner, and are not necessarily limited to those having all of the configurations described. Moreover, each processing unit of the data selection processing system (plant diagnosis system 1) shown in FIG. 1 is generally configured by installing software (programs) for executing each processing function as the data selection processing system in a computer device.

[0060] FIG. 13 shows an example of an outline of the hardware configuration of a computer device that functions as this data selection processing system.

[0061] The data selection processing system, which is constituted by a computer device, includes a CPU (Central Processing Unit) 1a, a main memory unit 1b, a non-volatile storage 1c, a network interface 1e, an input / output unit 1f, and a display unit 1g, which are all connected to a bus.

[0062] The CPU 1a is an arithmetic processing unit that reads out and executes program code of software that realizes the functions performed by the data selection processing system from the main storage unit 1b or the non-volatile storage 1c.

[0063] The CPU 1a reads out program codes from the main memory 1b or the non-volatile storage 1c and executes arithmetic processing in the work area of ​​the main memory 1b, thereby configuring various processing function units in the main memory 1b. For example, the main memory 1b is configured with a sensor data acquisition unit 11, an analysis group creation unit 13, an analysis unit 15, and an input / output control unit 18 shown in FIG.

[0064] The non-volatile storage 1c may be, for example, a large-capacity information storage medium such as a hard disk drive (HDD), a solid state drive (SSD), a memory card, etc. The non-volatile storage 1c stores software that realizes the functions of the data selection processing system, data obtained by executing the program, and information as each of the databases 12, 14, 16, and 17.

[0065] The network interface 1e is, for example, a network interface card (NIC) and is used for transmitting and receiving data to and from other devices. The input / output unit 1f takes in sensor data from the plant 3 and so on. The display unit 1g displays a screen such as that shown in Fig. 8. In the example of Fig. 1, the input / output device 2 performs the display, and the display unit 1g may not be provided in a computer device as a data selection processing system.

[0066] In addition, in the configuration diagram shown in Figure 1, only the control lines and information lines that are considered necessary for explanation are shown, and not all control lines and information lines in the product are necessarily shown. In reality, it can be considered that almost all components are connected to each other. Furthermore, when the system described in each embodiment is configured with an information processing device such as a computer, the programs that realize each processing function may be prepared in non-volatile storage or memory within the computer device, or may be stored on a recording medium such as an external memory, an IC card, an SD card, or an optical disk, and transferred. [Explanation of symbols]

[0067] 1...Plant diagnosis system (data selection processing system), 1a...CPU, 1b...Main memory unit, 1c...Non-volatile storage, 1e...Network interface, 1f...Input / output unit, 1g...Display unit, 2...Input / output device, 3...Plant, 11...Sensor data import unit, 12...Sensor value database, 13...Analysis group creation unit, 14...Analysis group database, 15...Analysis unit, 16...Analysis result database, 17...Tag number database, 18...Input / output control unit, 101-104, 104B...Window, 105...Button, 106...Result display field, 107...Sensor number display field, 108...Used sensor number display field, 109...Correlation coefficient display field, 201...Tab, 202...List, 203...Input sensor name display field, 204...Analysis result display field

Claims

1. a sensor data acquisition unit that acquires detection data from a plurality of sensors; an analysis unit that calculates a correlation coefficient of the detection data of the plurality of sensors that are acquired by the sensor data acquisition unit; an analysis group creation unit that defines, as a group, a set of sensors in which the correlation coefficient obtained by the analysis unit is higher than or equal to a threshold value set between the sensors; the analysis unit creates a graph in which each sensor is set as a node and edges are set between sensors whose correlation coefficients are equal to or greater than a threshold; A set of sensors that are highly correlated with each other is determined based on a clique or clique community. Data selection and processing system.

2. When the number of sensors belonging to the groups created by the analysis group creation unit exceeds a specified value, the number of sensors belonging to each group is adjusted by increasing the threshold value of the correlation coefficient.

2. The data selection and processing system according to claim 1.

3. When the number of sensors used by the analysis unit for analysis exceeds a specified value, the number of sensors used for analysis is adjusted by increasing the threshold value of the correlation coefficient.

3. The data selection and processing system according to claim 1 or 2.

4. When the number of groups created by the analysis group creation unit exceeds a specified value, the threshold value of the correlation coefficient is decreased, and when the number of groups created falls below the specified value, the threshold value of the correlation coefficient is increased, thereby adjusting the number of groups.

2. The data selection and processing system according to claim 1.

5. If there is a clique or clique community whose number of nodes exceeds a prescribed value, the correlation coefficient threshold is increased for the sensors belonging to the clique or clique community, and the graph is redrawn. In the reconstructed graph, if a new clique or clique community with a number of nodes within a specified range is generated, the process of defining the sensors belonging to this clique or clique community as a group is repeated, so that the number of sensors belonging to each group is within the specified range.

2. The data selection and processing system according to claim 1.

6. When the number of sensors used exceeds the specified value, the correlation coefficient threshold is increased, the graph is re-created, and the number of sensors used is adjusted.

6. A data selection and processing system according to claim 1 or 5.

7. When the number of groups created exceeds the specified value, the correlation coefficient threshold is decreased and the graph is re-created. When the number of groups created falls below the specified value, the correlation coefficient threshold is increased and the graph is re-created to adjust the number of groups.

6. A data selection and processing system according to claim 1 or 5.

8. a sensor data acquisition process for acquiring detection data from a plurality of sensors; an analysis process for calculating a correlation coefficient of the detection data of the plurality of sensors acquired by the sensor data acquisition process; and an analysis group creation process for defining, as a group, a set of sensors whose correlation coefficients obtained by the analysis process are higher than or equal to a threshold value set between the sensors, In the analysis process, each sensor is set as a node, and a graph is created in which edges are set between sensors whose correlation coefficient is equal to or greater than a threshold value; A set of sensors with high correlation with each other is determined based on a clique or clique community. Data selection processing method.

Citation Information

Patent Citations

  • Abnormality detection device and abnormality detection method

    JP2020201890A

  • Method and system for executing root cause analysis automated for abnormal event in high dimensional sensor data

    JP2021002335A

  • Failure sign diagnosis system and method for the same

    JP2021028751A

  • Abnormality analysis method, program, and system

    WO2018104985A1