A sequence similarity analysis method integrating alarm events and trend events
By integrating the sequence similarity analysis method of alarm events and trend events in industrial processes, the improved Smith-Watman algorithm and hierarchical clustering method are used to solve the problem of high similarity between different fault alarm sequences in complex industrial processes, and more accurate fault distinction and real-time monitoring are achieved.
Patent Information
- Application Number
- CN202311055633.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-21
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2043-08-21
AI Technical Summary
In complex industrial processes, due to the complexity of equipment and the strong coupling between internal variables, the alarm sequences of different faults are relatively similar and difficult to distinguish.
A sequence similarity analysis method that combines alarm events and trend events is adopted to generate multi-valued alarm signals by setting alarm thresholds to extract alarm events; qualitative trend analysis is used to extract trend characteristics of process signals, generate qualitative trend signals and extract trend events; alarm events and trend events are fused, and sequence comparison is used to use the improved Smith-Watman algorithm, and faults are distinguished by hierarchical clustering method.
It realizes better fault distinction when the alarm sequences of different faults are highly similar, and meets the decision-making needs of real-time monitoring and diagnosis of actual industrial industries.
Smart Images

Figure CN117113178B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of industrial process fault monitoring and diagnosis, and in particular to a sequence similarity analysis method integrating alarm events and trend events. Background Art
[0002] With the upgrading of industrial system equipment, the connection between equipment is becoming more and more complex, making the complexity of control systems increasingly increasing, making the decision-making task in the monitoring process difficult. In order to avoid unexpected downtime, it is necessary to use the collected process data information for comparative analysis. In the alarm system, when the data of the process signal exceeds the set alarm threshold, it is considered that an alarm has occurred. The occurrence of a fault usually produces a series of regular alarms or trend changes. The alarm condition is usually accompanied by abnormal changes in the process signal as a precursor. These changes are reflected not only in the amplitude of the process signal, but also in its direction. Therefore, both amplitude changes and trend changes are non-negligible features of the process signal. Since the same fault may cause similar alarms or trend change sequences, sequence similarity analysis can be introduced to effectively distinguish different faults.
[0003] The existing event sequence similarity analysis methods for complex industrial processes are mainly aimed at alarm sequences, which reflect the amplitude variation characteristics of process data. At present, the research content of alarm sequence similarity analysis methods mainly includes similarity analysis of paired sequences, multiple sequence alignment, cross-device sequence alignment, online alignment of alarm sequences and prediction methods. The above methods all generate alarm signals by setting alarm thresholds for process signals, and then extract alarm events in chronological order to obtain alarm sequences; finally, they are compared with the historical alarm sequence database for similarity. The alarm sequences of the same fault often have a high degree of similarity, and then the faults are distinguished by clustering or classification methods.
[0004] However, due to the complexity of the equipment and the strong coupling between its internal variables, the same variable may lead to similar alarm signals under different faults, which will lead to the situation that the alarm sequences of different faults are highly similar and difficult to distinguish. Therefore, in order to obtain a better fault differentiation effect, a sequence similarity analysis method that integrates alarm events and trend events is designed to improve the decision-making efficiency and productivity during operation. Summary of the invention
[0005] In order to solve the above problems, the present invention provides a sequence similarity analysis method that integrates alarm events and trend events to achieve fault differentiation. The original data collected from the actual industrial process is preprocessed to obtain high-quality analysis data; the alarm threshold is set to generate a multi-value alarm signal and then extract the alarm event; the trend characteristics of the preprocessed process signal are extracted using a qualitative trend analysis method, and the trend event is extracted after the qualitative trend signal is generated; the alarm event and the trend event are fused to obtain multi-feature events and sequences; the improved Smith-Waterman algorithm (SW) is used to perform sequence comparison to obtain the similarity between different sequences; the hierarchical clustering method is used to cluster the feature fusion sequences of different faults for fault differentiation.
[0006] The method of the present invention comprises the following steps:
[0007] S1. Extract the alarm sequence from the amplitude of the equipment process signal;
[0008] S2. Extract trend sequence from direction for equipment process signal;
[0009] S3, fusing the alarm sequence and the trend sequence to obtain a fused sequence;
[0010] S4. Perform similarity calculation and cluster analysis on the fused sequence to obtain different fault categories.
[0011] The beneficial effect provided by the present invention is that better fault differentiation can be achieved when the alarm sequences of different faults are highly similar and difficult to distinguish, and the decision-making needs of actual industrial real-time monitoring and diagnosis can be met. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 It is a schematic flow chart of the method of the present invention;
[0013] Figure 2 is the similarity heat map of alarm sequences;
[0014] Figure 3 is the similarity heat map of fusion sequences;
[0015] Figure 4 is the hierarchical clustering diagram of the alarm sequence;
[0016] Figure 5 is the hierarchical clustering graph of the fused sequence;
[0017] Figure 6 Is the alarm sequence and The local optimal matching graph of ;
[0018] Figure 7 Is a fusion sequence and The local optimal matching graph of . DETAILED DESCRIPTION
[0019] To make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be further described below with reference to the accompanying drawings.
[0020] Please refer to Figure 1 , Figure 1 It is a schematic diagram of the system structure of the present invention.
[0021] The present invention provides a sequence similarity analysis method for integrating alarm events and trend events, comprising:
[0022] S1. Extract the alarm sequence from the amplitude of the equipment process signal;
[0023] Step S1 is specifically as follows:
[0024] S11. Use wavelet denoising to denoise the process signal and obtain time series data x(t) = [x(1), x(2), ..., x(N)]; Based on the time series data, obtain the multi-alarm state x a (t); The multiple alarm states include: a high alarm state, a low alarm state and a normal state;
[0025] S12, from the multi-alarm state, using the ε-sampling closing delay timer to eliminate the chattering alarm and obtain the state after the chattering alarm is eliminated;
[0026] S13, extracting the alarm events of the single process variables from the state after the chattering alarm is eliminated, and then arranging and integrating all the alarm events in order of events to form an alarm sequence.
[0027] As an example, factories generally collect, extract and analyze data of many process variables in the equipment to monitor the system operation status as safely as possible. The monitoring equipment compares the current value of the process data with the set alarm threshold. When the data is greater than or less than the alarm threshold, it is considered an alarm.
[0028] A typical binary alarm signal is: 1 for alarm occurrence and 0 for normal. There are two forms of binary alarm signals: one is to regard the entire alarm period as an alarm occurrence and the rest as normal; the other is to regard only the alarm occurrence time as an alarm occurrence and the rest as normal. The generation of alarm signals is often accompanied by the appearance of chattering alarms, which may cause the extracted alarm sequence to be inaccurate, thereby affecting the similarity results of sequence comparison. To avoid large errors in similarity analysis, the chattering alarm should be removed first. After eliminating interference alarms and obtaining critical alarms, the alarm events and alarm sequences are extracted.
[0029] The steps are as follows:
[0030] (1) Generation of alarm signal: The process data is denoised by wavelet to obtain time series data x(t) = [x(1), x(2), …, x(N)], and a certain number of samples n of normal operation data are taken, and the mean value is
[0031]
[0032] Using the 3δ principle to determine the alarm threshold (T U is the upper limit of the alarm threshold, T L is the lower limit of the alarm threshold). Let λ=3, T U 、T L The calculation formula is as follows
[0033]
[0034] The present invention adopts multiple alarm states: high alarm (1), low alarm (-1), normal (0), which are defined as follows
[0035]
[0036] (2) Elimination of interference alarms: Common chatter elimination techniques include filters, delay timers, and dead zones, among which the delay timer can be directly applied to the alarm signal. The closed delay timer using ε sampling can effectively eliminate chatter alarms, which is defined as follows:
[0037]
[0038] Wherein, k = t-ε, t-ε+1, ..., t-1, ε is the number of samples of the closed delay device.
[0039] (3) Alarm event and sequence extraction: First, extract the alarm events of a single process variable separately, and then arrange and integrate all alarm events in chronological order to form an alarm sequence.
[0040] In the present invention, an alarm event is defined as a variable that has a certain alarm category at a certain moment, and a i =(t i ,X a ) to represent, where is the number of alarm events of a single process variable during the fault period, t i is the time when the alarm occurs, X a is the alarm category. Assuming that the number of process variables monitored by the alarm system is m, the total number of alarm events generated by all variables is The alarm sequence is defined as:
[0041]
[0042] Table 1 is an example of an alarm sequence for an industrial fault: the first column is the variable label, such as Tag.12, which means the 12th variable monitored; the second column is the alarm category, such as HI for high alarm and LO for low alarm; the third column is the time when the alarm event occurs.
[0043] Table 1 Alarm sequence examples for industrial faults
[0044]
[0045] S2. Extract trend sequence from direction for equipment process signal;
[0046] Step S2 is specifically as follows:
[0047] S21. Use the mean and variance method to test the stationarity of the process signal and obtain a non-stationary time series;
[0048] S22, for non-stationary time series, use PLR to extract trends and obtain trend segments;
[0049] S23. According to the trend segmentation, the trend is qualitatively classified into five trends: rapid rise, rise, fall, rapid fall and stable;
[0050] S24, respectively extracting trend events of a single process variable during the fault period, then integrating all trend events in chronological order to form a trend sequence, and merging the trend sequence according to trend segments.
[0051] As an embodiment, modern industrial plants usually store rich and massive data, but less effective information can be extracted from them. The alarm signal is extracted from the amplitude characteristics of the process signal. Therefore, it is also necessary to extract the characteristics of the process signal in direction (i.e., trend changes). Piece-wise linear representation (PLR) is one of the commonly used and effective trend extraction techniques. Its principle is to separate a long time series into multiple data segments, and each data segment is represented by a straight line. This processing can extract the potential trend changes of the process signal. Qualitative Trend Analysis (QTA) is an effective tool for analyzing the trend changes of process signals. Its purpose is to extract and qualitatively express the trend changes of process signals, so as to facilitate comparative analysis.
[0052] Since the data of different process variables will show different trend changes during the fault period, some process variables will show obvious changes, while others will be relatively stable. Therefore, it is necessary to perform a stability test on the data of the process variables before trend extraction. If the test result is non-stationary, trend extraction is performed and qualitative expression is made. The qualitative trend changes of two or more consecutive data segments obtained by PLR may be the same. However, the extraction of trend events is aimed at the time points when trend changes occur between data segments. Therefore, merging two or more consecutive trend events with the same trend changes is a good solution. Subsequently, the trend sequence extraction process is similar to the alarm sequence extraction.
[0053] The specific steps are as follows:
[0054] (1) Stationarity test: Use the mean and variance method to test the stationarity of the data. For a time series x(t) = [x(1), x(2), ..., x(N)], where N is the length of the time series, divide it into L consecutive segments and use [P1, P2, ..., P L ] represents. Let v i is the mean of each continuous segment, i=1,2,...,L, The significance level based on the mean stationarity test is
[0055]
[0056] Among them, σ1 is the average value For i=1,2,...,L, if there exists or Then the time series is non-stationary, otherwise it is considered stationary.
[0057] (2) Trend extraction: PLR is used to extract trends from non-stationary time series. Divide x(t) into M data segments connected end to end, and the mth data segment will be represented as [x(1:t m ),...,x(1:t m+1 -1)], where t m and t m+1 -1 is the first and last data points of the segment. Assume that a linear regression model is used for each data segment and the approximation is
[0058]
[0059] Among them, t∈[t m ,t m+1 -1], and are the slope and intercept, respectively, estimated by the least squares method and minimized by the fitting error Next, we need to determine a specific number of segments M. If C m is the parallelogram space of confidence intervals for the estimates, and G m is the convex hull composed of these sample data, then it means C k The index η of the percentage of overlapping area in m Then
[0060]
[0061] The index function of the number of segments M is expressed as
[0062]
[0063] This is essentially the exponential η m The weighted average of m = 1, 2, ..., M, the final number of segments M is
[0064]
[0065] (3) Qualitative expression of trends: After the time series is extracted through PLR, the trend changes of rising, falling, and stable are often used to qualitatively express it. From the above-mentioned assumed linear regression model, we can know that For each data segment [x(t m ),…,x(t m+1 -1)], the mean of the M data segments is based on The standard deviation of
[0066]
[0067] Since the rate of rise or fall can be fast or slow, the present invention will characterize the trend change into five situations: rapid rise (2), rise (1), fall (-1), rapid fall (-2), and stability (0). The trend signal is defined as follows:
[0068]
[0069] in, is a stable interval, and are the rising and falling slope thresholds, respectively, To distinguish between rising and fast rising slope thresholds, is the slope threshold for distinguishing between a decline and a rapid decline. In the present invention, λ1 and λ2 are set to 1 / 4 and 1 respectively.
[0070] (4) Trend sequence extraction: first extract the trend events of a single process variable during the fault period, and then integrate all trend events in chronological order to form a trend sequence. In the present invention, a trend event is defined as a certain trend change of a variable at a certain moment, and q i =(t i ,X t ) to represent, where is the number of trend events of a single process variable during the fault period, t i When the trend changes, X t is the trend change category. Assuming that the number of process variables monitored by the alarm system is m, the total number of alarm events generated by all variables is The trend series is defined as:
[0071]
[0072] Table 2 is an example of a trend sequence of industrial failures based on Table 1: the first column is the variable label, such as Tag.22, which means the 22nd variable monitored; the second column is the trend category, such as UUP for rapid increase, UP for increase, DDN for rapid decrease, DN for decrease, and ST for stability; the third column is the time when the trend event occurs.
[0073] Table 2 Example of trend sequence of industrial failures
[0074]
[0075] (5) Trend segment merging: Extract trend events from a single process variable and arrange them in chronological order to obtain S1 = <(t1,1), (t2,1), (t3,2), (t4,-2), (t5,-1)>. Since the trend changes at both t1 and t2 are upward, the event (t2,1) is deleted and the final sequence is updated to S1 = <(t1,1), (t2′,2), (t3′,-2), (t4′,-1)>, where t2′ = t3, t3′ = t4, and t4′ = t5.
[0076] S3, fusing the alarm sequence and the trend sequence to obtain a fused sequence;
[0077] Step S3 is as follows:
[0078] S31. Construct feature event f i =(t f ,X f ), the characteristic event refers to a change in a certain characteristic at a certain moment, including an alarm state change and a trend change; where i = 1, 2, ... N f , Nf is the total number of characteristic events contained in the fusion sequence, t f is the feature occurrence time, X f is the feature category;
[0079] S32, the fusion sequence is:
[0080] It should be noted that in the actual industrial production process, the alarm sequence is composed of a certain pattern or a series of alarm signals sorted in chronological order. Sequence similarity analysis is often used to reflect the similarity between paired sequences or multiple sequences. Sequences with high similarity are often caused by the same fault or abnormality. By comparing the similarity levels between different alarm sequences, it helps the operator to analyze these alarm sequence patterns, so as to determine the most likely cause of the alarm and formulate corresponding strategies. The above two main steps S1 and S2 are mainly to obtain alarm events and trend events during the fault period. The ultimate goal of the present invention is to perform similarity calculation and cluster analysis on the fused sequence. Therefore, it is necessary to fuse the alarm events and trend events to obtain the fused sequence.
[0081] In the present invention, a characteristic event is defined as a characteristic change of a variable at a certain moment, including alarm and trend change, and is represented by f i =(t f ,X f ), where i = 1, 2, ... N f , N f is the total number of characteristic events contained in the fusion sequence (i.e., the total number of alarm events and trend events of all process variables), t f is the feature occurrence time, X f is the feature category. The fusion sequence is more representative and different from the alarm sequence, and is defined as follows:
[0082]
[0083] The fusion sequence shown in Table 3 can be obtained by fusing Table 1 and Table 2. In the alarm sequence and trend sequence, there are cases where the variable labels are the same but represent different meanings. For example, Tag.12 in Table 1 means: when fault type 1 occurs, the 12th process variable signal is monitored to have a high alarm; while in Table 2, it means: when the fault occurs, the 12th process variable signal is monitored to have a rapid rising trend change. Therefore, variable labels need to be distinguished, such as Tag.12A means that the 12th process variable has an alarm event, and Tag.12T means that the 12th process variable has a trend event.
[0084] Table 3. Examples of fusion sequences of industrial faults
[0085]
[0086] S4. Perform similarity calculation and cluster analysis on the fused sequence to obtain different fault categories.
[0087] Step S4 is specifically as follows:
[0088] S41, using the improved SW method to compare multiple fusion sequences to obtain a similarity matrix;
[0089] Step S41 is specifically as follows:
[0090] S411, calculate fusion sequence f i 1 and The element in the i-th row and j-th column of the matching score matrix H is:
[0091]
[0092] Among them, α and β are the matching and mismatching scores respectively, Among them, |t m -t i | represents the time interval between the mth alarm and the nearest alarm identifier k on the time axis, k = 1, 2, ..., Z, Z is the total number of alarm types; σ3 represents the bandwidth of the time interval;
[0093] S412, find the maximum value in the matching score matrix H, and backtrack along the score path to obtain the local optimal match of sequences S1 and S2;
[0094] S413. Using Similarity Index Measure the similarity between the best matching sequences S1 and S2:
[0095]
[0096] Where m1 and m2 are the lengths of sequences S1 and S2 respectively, and o(S1, S2) is the maximum value in H;
[0097] S414, performing mutual comparison on multiple groups of fusion sequences using steps S411 to S413 to obtain multiple similarities and form a similarity matrix.
[0098] S42. Based on the similarity matrix, hierarchical clustering method is used to perform clustering to obtain different categories.
[0099] It should be noted that, at present, sequence alignment algorithms are divided into two categories: global sequence alignment algorithms and local sequence alignment algorithms. Local sequence alignment methods are usually used in the industrial field. The present invention adopts an improved SW algorithm, which essentially uses timestamp information to blur the order of variables in the alarm sequence when the alarm occurrence time is close to each other, making it no longer sensitive to the alarm sequence. The advantage of such processing is that more similar parts between sequences can be found when a large number of alarms occur almost at the same time. The similarity matrix between sequences is obtained by sequence alignment, and a hierarchical clustering method is used to cluster them, and the clustering tree is observed and the purity evaluation index is used to measure the clustering effect.
[0100] Taking the traditional Simth-Waterman (SW) algorithm as an example, for the sequence Among them, n1 and n2 are the lengths of the two sequences, and the i-th row and j-th column in the matching score matrix H are The calculation is
[0101]
[0102] Where i = 1, 2, ..., n1, j = 1, 2, ..., n2, δ is the vacancy penalty, For element f i 1 With elements The matching score is defined as
[0103]
[0104] Where α and β are the match and mismatch scores, respectively. Usually, α>0, β<0, δ<0, and α>max(|β|,|δ|). In order to prefer indirect alignment rather than mismatch, the parameter should be set to β<2δ<0.
[0105] The principle of the improved SW algorithm is as follows: Assume that there are Z types of alarms, and the time weight The determination is based on the following proportional Gaussian function, namely
[0106]
[0107] Among them, |t m -t i | represents the time interval between the mth alarm and the nearest alarm identifier k on the time axis, k = 1, 2, …, Z, and σ3 represents the bandwidth of the time interval.
[0108] For each item in H The calculation is
[0109]
[0110] After obtaining H, backtrack it, that is, start from the item with the largest score, backtrack according to the score path, and find the local optimal match.
[0111] Finally, through a similarity index To measure the similarity between sequences.
[0112] The present invention adopts Among them, S1 and S2 are different sequences, m1 and m2 are their lengths respectively, and o(S1, S2) is the maximum value in H.
[0113] The hierarchical clustering method in step S42 is specifically as follows:
[0114] ① Classify each object into one class, and get N classes in total, each class contains only one object. The distance between classes is the distance between the objects they contain;
[0115] ② Find the two closest classes and merge them into one class, so the total number of classes is reduced by one;
[0116] ③ Recalculate the distance between the new class and all old classes;
[0117] ④ Repeat the above two steps until they are finally merged into one class;
[0118] ⑤According to the different distances selected, the hierarchical clustering method can be divided into single linkage, full linkage and equal linkage clustering methods.
[0119] In order to evaluate the performance of hierarchical clustering, the present invention adopts a density-based purity index evaluation index. The overall idea of clustering purity is to divide the number of samples that are correctly clustered by the total number of samples, which is often called the accuracy of clustering. For the results after clustering, the true category corresponding to each cluster is unknown, so it is necessary to take the maximum value in each case. The calculation formula of purity is defined as follows:
[0120]
[0121] Where N is the total number of samples, c p represents the number of samples in the pth cluster after clustering, r o represents the number of true samples in the pth cluster, and P u ∈[0,1], the closer it is to 1, the better the clustering result.
[0122] Finally, as an embodiment, this embodiment designs a simulation experiment based on the TEP model to verify the effectiveness of the method.
[0123] (1) Data Collection
[0124] The data used is from the TEP model, which has been widely regarded as a benchmark for control and monitoring research. The TEP simulation model is a realistic simulation program for a chemical plant. It defines 20 fault types (the Disturbances module can set different faults) and allows 52 measurements (41 process variables and 11 operating variables). In the TEP simulation model structure diagram, these 20 fault types and their addition times can be pre-set by the simulation model, such as 0 in normal working conditions, fault 1 is the first digit set to 1, and the others are 0; faults are added in sequence, fault 20 is the twentieth digit set to 1, and the others are 0. There are 15 known faults in the TE process and 5 unknown faults that have not been determined. Among them, faults 1 to 7 are faults related to step changes, faults 8 to 12 are faults related to increased variance, fault 13 is caused by slow drift of the reaction rate in the reactor, which is more suitable for fault prediction, and faults 14 and 15 are caused by valve failure.
[0125] In this example, the sampling period is set to 168h and the sampling frequency is set to 3min / time. By adjusting the Disturbances module to add different fault types at different time points, 6 fault types are selected for testing, namely Fault 1 (F1), Fault 2 (F2), Fault 6 (F6), Fault 7 (F7), Fault 8 (F8) and Fault 20 (F20). Four sets of process data are generated under the same fault, and a total of 24 sets of data are generated for the six faults, among which two sets are for the fault duration of 1h and 2h respectively. The details are shown in Table 4.
[0126] Table 4 Fault data record
[0127]
[0128] (2) Obtaining alarm sequence and fusion sequence
[0129] The first 22 process variables are selected. First, the alarm threshold is set for the process signal to obtain the alarm signal, and the interference alarm is eliminated by using a closed delayer. At the same time, the process signal is tested for stationarity, and the trend is extracted and qualitatively expressed by PLR for non-stationary models; then, the alarm events and alarm sequences are extracted, and at the same time, the same trend segments are merged to extract trend events and trend sequences; finally, the alarm events and trend events are fused to obtain a new feature sequence, namely the fused sequence.
[0130] (3) Similarity calculation based on sequence alignment
[0131] The improved SW algorithm is used to compare the alarm sequences and fusion sequences of different faults. First, the appropriate parameters are adjusted to make the gap penalty smaller while satisfying the gap alignment; then, the score path is recorded to trace back to find the optimal matching result; finally, the similarity index is calculated. To obtain the similarity matrix between sequences. Figure 2 and Figure 3 These are the similarity heat maps of the alarm sequence and the fusion sequence, respectively. The darker the color, the higher the similarity. In the alarm sequence similarity heat map, most of the colors are quite dark and difficult to distinguish. There are more than 4 dark blocks in the same row or column. In contrast, the color of the fusion sequence similarity heat map is weakened overall, and the color distinction is more intuitive.
[0132] (4) Event sequence cluster analysis
[0133] In order to observe the similarity results more intuitively, the similarity matrix of the alarm sequence and the fusion sequence obtained in the previous step is hierarchically clustered to obtain a hierarchical clustering tree to observe the fault differentiation results. Figure 4 and Figure 5 is the hierarchical clustering diagram of the alarm sequence and the fusion sequence. Tables 5 and 6 are the clustering result tables of the alarm sequence and the fusion sequence respectively.
[0134] Table 5 Clustering results of alarm sequences
[0135]
[0136] Indicates the i a Alarm sequence generated by group data, i a =1,2,…,24. It can be found that in C1 All are generated under F20, but It is generated under F7. Therefore, according to the definition of cluster purity, the total number of samples of C1 is 5, and the maximum number of true samples is 4. Similarly, the total number of samples of C2 is 3, and the maximum number of real samples is 3 The total number of samples of C3 is 6, and the maximum number of true samples is 4 The total number of samples of C4 is 4, and the maximum number of true samples is 4 The total number of samples of C5 is 6, and the maximum number of true samples is 4 Purity calculation under alarm sequence is:
[0137] Table 6 Clustering results of fusion sequences
[0138]
[0139] Indicates the i f Alarm sequence generated by group data, i f =1,2,…,24. Similarly, according to the definition of cluster purity, we can find that the total number of samples of C1 is 4, and the maximum number of true samples is 4. The total number of samples of C2 is 3, and the maximum number of true samples is 3 The total number of samples of C3 is 6, and the maximum number of true samples is 3 The total number of samples of C4 is 3, and the maximum number of true samples is 3 The total number of samples of C5 is 4, and the maximum number of true samples is 4 ; The total number of samples of C6 is 2, and the maximum number of true samples is 1 ( or ); The total number of samples of C7 is 4, and the maximum number of true samples is 4 Purity calculation under alarm sequence is:
[0140] like Figure 4 As shown, there are cases where the alarm sequences of different faults are similar and cannot be distinguished, for example, Not with and Classified as the same category, but Figure 5 middle, and and The clustering is correct. The reasons can be analyzed through the specific results of sequence matching. Figure 6 for and The local optimal matching result between Figure 7 for and The local optimal matching results between , where the samples are arranged in sampling order, from 1 to 3361.
[0141] As can be seen from the figure, and The number of matching events between and The sequence length is longer than the alarm sequence, which will cause the similarity to decrease. When compared with the fused sequences of different faults, many trend events cannot be matched, resulting in a significant decrease in similarity. Therefore, although the similarity between fused sequences decreases overall, the similarity between fused sequences of the same fault is higher than the similarity between other faults, so they are clustered into the same cluster, achieving accurate fault differentiation.
[0142] The above hierarchical clustering results and sequence alignment results show that similarity analysis based on fused sequences has better fault differentiation ability than similarity analysis based on alarm sequences alone. Therefore, it can be considered that the proposed sequence similarity analysis method that integrates alarm events and trend events is very effective in distinguishing different fault sequences.
[0143] The beneficial effect of the present invention is that better fault differentiation can be achieved when the alarm sequences of different faults are highly similar and difficult to distinguish, and the decision-making needs of actual industrial real-time monitoring and diagnosis can be met.
[0144] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A sequence similarity analysis method for integrating alarm events and trend events, characterized in that: The following steps are involved: S1. Extract the alarm sequence from the amplitude of the equipment process signal; Step S1 is specifically as follows: S11, using wavelet denoising to process signal, to obtain time series data x(t)=[x(1),x(2),…,x(N)]; According to the time series data, get the multi-alarm status x a (t); The multiple alarm states include: a high alarm state, a low alarm state and a normal state; S12, from the multi-alarm state, using the ε-sampling closing delay timer to eliminate the chattering alarm and obtain the state after the chattering alarm is eliminated; S13, extracting the alarm events of the individual process variables from the state after the chattering alarm is eliminated, and then arranging and integrating all the alarm events in order of events to form an alarm sequence; S2. Extract trend sequence from direction for equipment process signal; Step S2 is specifically as follows: S21. Use the mean and variance method to test the stationarity of the process signal and obtain a non-stationary time series; S22, for non-stationary time series, trend extraction is performed using PLR piecewise linear representation to obtain trend segmentation; S23. According to the trend segmentation, the trend is qualitatively classified into five trends: rapid rise, rise, fall, rapid fall and stable; S24, extracting trend events of a single process variable during the fault period respectively, then integrating all trend events in chronological order to form a trend sequence, and merging the trend sequence according to trend segments; S3, fusing the alarm sequence and the trend sequence to obtain a fused sequence; Step S3 is as follows: S31. Construct feature event f i =(t f ,X f ), the characteristic event refers to a change in a certain characteristic at a certain moment, including an alarm state change and a trend change; where i = 1, 2, ... N f , N f is the total number of characteristic events contained in the fusion sequence, t f is the feature occurrence time, X f is the feature category; S32, the fusion sequence is: S4. Perform similarity calculation and cluster analysis on the fused sequence to obtain different fault categories.
2. A method for analyzing sequence similarity of fused alarm events and trend events as claimed in claim 1, characterized in that: Step S4 is specifically as follows: S41, using the improved SW method to compare multiple fusion sequences to obtain a similarity matrix; S42. Based on the similarity matrix, hierarchical clustering method is used to perform clustering to obtain different categories.
3. A method for analyzing sequence similarity of fused alarm events and trend events as claimed in claim 2, characterized in that: After obtaining different categories, a pure index evaluation index based on density is also used to evaluate the clustering results.
4. The method for analyzing sequence similarity of fused alarm events and trend events according to claim 1, characterized in that: The alarm sequence is: Among them, a i =(t i ,X a ) is an alarm event, is the number of alarm events of a single process variable during the fault period, t i X is the time when the alarm changes. a is the alarm category; m is the number of process variables monitored by the system.
5. The method for analyzing sequence similarity of fused alarm events and trend events according to claim 1, characterized in that: The trend sequence is: Among them, q i =(t i ,X t ) is a trend event, is the number of trend events of a single process variable during the fault period, t i When the trend changes, X a is the trend category; m is the number of process variables monitored by the system.
6. A method for analyzing sequence similarity of fused alarm events and trend events as claimed in claim 2, characterized in that: Step S41 is specifically as follows: S411, calculate fusion sequence and The element in the i-th row and j-th column of the matching score matrix H is: Among them, α and β are the matching and mismatching scores respectively, Among them, |t m -t i | represents the time interval between the mth alarm and the nearest alarm identifier k on the time axis, k = 1, 2, ..., Z, Z is the total number of alarm types; σ3 represents the bandwidth of the time interval; S412, find the maximum value in the matching score matrix H, and backtrack along the score path to obtain the local optimal match of sequences S1 and S2; S413. Using Similarity Index Measure the similarity between the best matching sequences S1 and S2: Where m1 and m2 are the lengths of sequences S1 and S2 respectively, and o(S1, S2) is the maximum value in H; S414, performing mutual comparison on multiple groups of fusion sequences using steps S411 to S413 to obtain multiple similarities and form a similarity matrix.
7. The method for analyzing sequence similarity of fused alarm events and trend events according to claim 3, characterized in that: The pure index evaluation method based on density is as follows: Where N is the total number of samples, c p represents the number of samples in the pth cluster after clustering, r o represents the number of true samples in the pth cluster, and P u ∈[0,1], the closer it is to 1, the better the clustering result.