Monitoring method, terminal and medium based on graph embedding and cross-time segment distribution matching

By constructing a graph embedding distribution matching model across time segments, the problem of poor monitoring effect in non-stationary industrial processes is solved, and a higher fault detection rate and data feature extraction are achieved.

CN120469325BActive Publication Date: 2025-09-09TIANJIN POLYTECHNIC UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510969148.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-09-09
Estimated Expiration
2045-07-15

AI Technical Summary

Technical Problem

Existing data-driven process monitoring methods are ineffective in non-stationary industrial processes and cannot effectively handle the problem of temporal covariate drift, resulting in low fault detection rates.

Method used

By acquiring data from non-stationary industrial processes, dividing the time segments with the largest differences, and constructing a graph embedding distribution matching model across time segments, global monitoring indicators are constructed using Bayesian inference and kernel density estimation to achieve data distribution matching and structure preservation in low-dimensional space, and perform online fault detection.

Benefits of technology

It improves the fault detection rate of non-stationary industrial processes, captures the distribution changes of data and maintains the local and global geometric structure of the data, and extracts more comprehensive data features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120469325B_ABST
    Figure CN120469325B_ABST
Patent Text Reader

Abstract

The present invention provides a monitoring method, terminal, and medium based on graph embedding distribution matching across time segments, wherein the method includes: obtaining non-stationary industrial process data under normal operating conditions, constructing a graph embedding distribution matching model across time segments; constructing two global monitoring indicators based on Bayesian inference, and using kernel density estimation to calculate the control limits of the two global monitoring indicators to obtain the control limits of the offline modeling stage; obtaining new sample data of the industrial process online, calculating its global monitoring indicator, comparing it with the control limits, and judging whether the industrial process has failed based on the comparison results. The monitoring method, terminal, and medium based on graph embedding distribution matching across time segments described in the present invention can capture the distribution changes of non-stationary data, and while matching the distributions of different time segments, can maintain the local and global geometric structures of the data, so that this monitoring method can significantly improve the fault detection rate and has better application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of process monitoring and fault diagnosis, and in particular relates to a monitoring method, terminal and medium based on graph embedding and cross-time segment distribution matching. Background Art

[0002] Modern industry is developing towards complexity and large-scale development, and early failures may lead to serious accidents. Therefore, real-time process monitoring is crucial to ensuring the safety and reliability of industrial systems. Data-driven process monitoring methods have attracted widespread attention in academia in recent years. Due to the rapid development of information technology and sensor measurement technology, the large amount of data generated in industrial processes can be efficiently collected and stored, which provides a solid foundation for the development of data-driven process monitoring methods. Compared with model-based methods, data-driven methods do not need to rely on a large amount of prior knowledge or precise mathematical models, which significantly lowers the application threshold and further promotes its widespread application in the industrial field.

[0003] Traditional data-driven process monitoring methods, such as principal component analysis and canonical correlation analysis, have been widely used in industrial process monitoring. However, these methods often rely on a key assumption: the monitored process is stationary, meaning that the data distribution remains constant over time. This assumption often fails to hold in real industrial scenarios. Due to factors such as equipment aging and environmental disturbances, industrial processes can exhibit non-stationarity. This non-stationarity can lead to temporal covariate drift, manifesting as temporal changes in the distribution of process variables. Therefore, traditional data-driven methods are not suitable for modeling non-stationary industrial processes. In recent years, research on non-stationary industrial process monitoring has also been conducted. Adaptive methods can modify monitoring models in real time and are capable of handling non-stationarity in industrial processes. Recursive and moving window strategies are common adaptive methods that have been widely studied. However, in practical applications, these adaptive methods increase computational burden and are not adaptable to slow-onset fault data, which seriously affects the effectiveness of non-stationary industrial process monitoring. Summary of the Invention

[0004] In view of this, the present invention aims to propose a monitoring method, terminal and medium based on graph embedding and cross-time segment distribution matching to solve the problem that the existing technology has poor monitoring effect when monitoring non-stationary industrial processes.

[0005] To achieve the above object, the technical solution of the present invention is achieved as follows:

[0006] In a first aspect, the present invention provides a monitoring method based on graph embedding and cross-time segment distribution matching, comprising:

[0007] Obtain non-stationary industrial process data under normal operating conditions and segment the data to obtain the time segments with the largest distribution differences;

[0008] Construct a graph embedding cross-time segment distribution matching model; wherein, the graph embedding cross-time segment distribution matching model is constructed by learning projection , the non-stationary industrial process data of different time segments are projected into a low-dimensional space, and the distribution matching of different time segments is achieved in the low-dimensional space, while the structure of the non-stationary industrial process data is maintained by constructing a joint objective function;

[0009] Two global monitoring indicators are constructed based on Bayesian inference, and the control limits of the two global monitoring indicators are calculated using kernel density estimation to obtain the control limits of the offline modeling stage. Among them, one global monitoring indicator is used to monitor data changes in the principal component space, and the other global monitoring indicator is used to monitor data changes in the residual space.

[0010] New sample data of the industrial process is obtained online, and two global monitoring indicators corresponding to the new sample data are calculated and compared with the control limits in the offline modeling stage to obtain a comparison result, and whether the industrial process has a fault is determined based on the comparison result.

[0011] Furthermore, the process of obtaining industrial process data under normal operating conditions and obtaining time segments with the largest distribution differences by segmenting the data may include:

[0012] Acquisition of non-stationary industrial process data under normal conditions ,in is the dimension of the sample data, is the number of samples, to Indicates 1 to samples;

[0013] The non-stationary industrial process data is normalized so that each column of variables has a mean of 0 and a variance of 1. The formula is as follows:

[0014] (1)

[0015] In the above formula express Any sample in is the mean vector of the variables, is the variance vector of the variables;

[0016] The objective function is used to segment the normalized non-stationary industrial process data. The formula is as follows:

[0017] (2)

[0018] In the above formula Used to calculate the number of cumulative summations, Indicates the number of time segments, the subscript in the above formula and Subscripts representing different time segments, and Indicates different time segment data, and there are differences in the distribution of the two. Represents a time segment The sample points, Represents a time segment The sample points, Represents a time segment The number of samples, Represents a time segment The number of samples, and It is a preset parameter used to avoid unnecessary solutions;

[0019] First, the normalized non-stationary industrial process data are divided into 2 to share( and ),get There are 1 to 2 split strategies. split points; then, under each segmentation strategy, the distribution difference in different cases is calculated using formula (2), and the segmentation strategy with the largest distribution difference is selected as the final solution.

[0020] Furthermore, constructing a graph embedding cross-time segment distribution matching model includes:

[0021] The maximum average difference is used to match the distribution between different time segments, and the first objective function is used to reduce the distribution difference of data in different time segments to reduce the non-stationarity of the data. The formula is as follows:

[0022] (3)

[0023] In the above formula and Represents time segments and The sample points, and Represents time segments and The optimization problem is expressed in the form of a trace, as follows:

[0024] (4)

[0025] (5)

[0026] (6)

[0027] (7)

[0028] In the above formula (4) , Indicates the number of time segments; Indicates the number of elements is The all-one vector of Indicates the number of elements is The all-one vector of 、 、 represents a fully unified matrix, and Equations (5), (6), and (7) represent the Matrix method;

[0029] The second objective function is constructed to maintain the local geometric structure of each time segment. The formula is as follows:

[0030] (8)

[0031] In the above formula It's a time segment The adjacency matrix of is as follows:

[0032] (9)

[0033] In the above formula is the thermonuclear parameter, express of The optimization problem is written in the form of traces, and the formula is as follows:

[0034] (10)

[0035] In the above formula Represents the data of the first time segment, Represents the data of the last time segment, represents the Laplacian graph, is a diagonal matrix, represents an all-zero matrix;

[0036] The third objective function is constructed to maintain the global geometric structure of each time segment. The formula is as follows:

[0037] (11)

[0038] In the above formula is the non-adjacency matrix of each time segment, and the formula is as follows:

[0039] (12)

[0040] In the above formula is the thermonuclear parameter, represents the geodesic distance between two sample points in the same time segment; where the optimization problem is expressed in the form of a trace, the formula is as follows:

[0041] (13)

[0042] In the above formula represents the Laplacian graph, is a diagonal matrix;

[0043] Combining the optimization problems corresponding to the first, second, and third objective functions, a graph embedding distribution matching model across time segments is constructed. The formula is as follows:

[0044] (14)

[0045] In the above formula is the weight coefficient of each item, and the joint objective function is transformed into the following formula:

[0046] (15)

[0047] The joint objective function is solved using the Lagrangian technique, and the optimization problem corresponding to the joint objective function is converted into a generalized eigenvalue decomposition to obtain the projection , the formula is as follows:

[0048] (16)

[0049] In the above formula yes The largest eigenvalues, where the number of selected eigenvalues ​​must be less than the sample dimension.

[0050] Furthermore, the two global monitoring indicators are constructed based on Bayesian inference, and the control limits of the two global monitoring indicators are calculated using kernel density estimation to obtain the control limits of the offline modeling stage, including:

[0051] Using projection Project each time segment into its own subspace to establish local models, and calculate two statistics belonging to the model samples under each local model. The formula is as follows:

[0052] (17)

[0053] (18)

[0054] In the above formula Statistics are used to monitor data changes in the principal element space. Statistics are used to monitor data changes in the residual space. is the low-dimensional representation of the sample, Represents a time segment covariance of the low-dimensional representation, represents the residual space;

[0055] Calculate the control limits using kernel density estimation, using Represents a time segment of Statistical control limits, and use Represents a time segment of Statistical control limits;

[0056] Bring each sample into The local model calculates statistics and constructs the The global monitoring indicator is composed of the local results, and the probability of each sample belonging to each local model is calculated. The formula is as follows:

[0057] (19)

[0058] In the above formula represents the marginal probability, and are prior probability and conditional probability respectively;

[0059] The prior probability Calculated by the following formula:

[0060] (20)

[0061] Secondly, the conditional probability of the sample under the two statistics is calculated by the following formula:

[0062] (twenty one)

[0063] (twenty two)

[0064] In the above formula express of Statistic conditional probability, express of Statistic conditional probability, express In the time segment model Next Statistics, express In the time segment model Next Statistics, Represents a time segment of Statistical control limits, Represents a time segment of Statistical control limits;

[0065] The two global monitoring indicators are calculated using the following weighted averages:

[0066] (twenty three)

[0067] (twenty four)

[0068] In the above formula express exist Statistics belong to time segments The probability of express exist Statistics belong to time segments probability;

[0069] The control limits of the two global monitoring indicators are calculated using kernel density estimation to obtain the control limits in the offline modeling stage.

[0070] Furthermore, the online acquisition of new sample data of the industrial process and the calculation of two global monitoring indicators corresponding to the new sample data are compared with the control limits of the offline modeling stage to obtain a comparison result, and judging whether a fault occurs in the industrial process based on the comparison result includes:

[0071] Collect new sample data , and normalized;

[0072] Calculate the statistics of the new sample data under each local model. The formula is as follows:

[0073] (25)

[0074] (26)

[0075] In the above formula Is a low-dimensional representation of the new sample data; where ;

[0076] Calculate the probability that the new sample data belongs to each local model. The formula is as follows:

[0077] ; (27)

[0078] Calculate the conditional probability of new sample data under two statistics. The formula is as follows:

[0079] (28)

[0080] (29)

[0081] In the above formula express of The statistic conditional probability, express of The conditional probability of a statistic, express In the time segment model Next Statistics, express In the time segment model Next Statistics;

[0082] Calculate the global monitoring indicators of new sample data using the following formula:

[0083] (30)

[0084] (31)

[0085] In the above formula express exist Statistics belong to time segments The probability of express exist Statistics belong to time segments probability;

[0086] Determine whether the global monitoring indicators of the new sample data exceed the control limit. If they exceed the control limit, it is fault data.

[0087] In a second aspect, the present invention further provides a monitoring device based on graph embedding and cross-time segment distribution matching, comprising:

[0088] The offline module is used to obtain non-stationary industrial process data under normal operating conditions and to obtain the time segments with the largest distribution differences by segmenting the data;

[0089] A construction module is used to construct a graph embedding cross-time segment distribution matching model; wherein, the graph embedding cross-time segment distribution matching model is constructed by learning projection , the non-stationary industrial process data of different time segments are projected into a low-dimensional space, and the distribution matching of different time segments is achieved in the low-dimensional space, while the structure of the non-stationary industrial process data is maintained by constructing a joint objective function;

[0090] The calculation module is used to construct two global monitoring indicators based on Bayesian inference and calculate the control limits of the two global monitoring indicators using kernel density estimation to obtain the control limits in the offline modeling stage. Among them, one global monitoring indicator is used to monitor data changes in the principal component space, and the other global monitoring indicator is used to monitor data changes in the residual space.

[0091] The comparison module is used to obtain new sample data of the industrial process online, calculate the two global monitoring indicators corresponding to the new sample data, compare them with the control limits of the offline modeling stage to obtain comparison results, and judge whether the industrial process has failed based on the comparison results.

[0092] In a third aspect, the present invention further provides a terminal, comprising:

[0093] one or more processors;

[0094] a storage device for storing one or more programs;

[0095] Sensors, used to collect data;

[0096] When the one or more programs are executed by the one or more processors, the one or more processors implement the monitoring method based on graph embedding and cross-time segment distribution matching as described above.

[0097] In a fourth aspect, the present invention further provides a storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to perform the monitoring method based on graph embedding and cross-time segment distribution matching as described above.

[0098] In a fifth aspect, an embodiment of the present invention further provides a computer product, including a computer program, which, when executed by a processor, implements the monitoring method based on graph embedding and cross-time segment distribution matching as described above.

[0099] Compared with the existing technology, the monitoring method, terminal, and medium based on graph embedding and cross-time segment distribution matching described in the present invention have the following advantages:

[0100] The monitoring method, terminal, and medium for cross-time segment distribution matching based on graph embedding described in the present invention can capture the distribution changes of non-stationary data. While matching the distribution of different time segments, it can maintain the local and global geometric structure of the data and characterize the intrinsic geometric information of the data, so that this monitoring method can extract more comprehensive data features and has better application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0101] The accompanying drawings, which constitute part of the present invention, are provided to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are provided to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings:

[0102] Figure 1 This is a flow chart of a monitoring method based on graph embedding and cross-time segment distribution matching according to the first embodiment of the present invention;

[0103] Figure 2 A schematic diagram of temporal covariate drift in a monitoring method based on graph embedding and cross-time segment distribution matching according to the first embodiment of the present invention;

[0104] Figure 3 Schematic diagram of a monitoring framework of a monitoring method based on graph embedding and cross-time segment distribution matching according to the first embodiment of the present invention;

[0105] Figure 4 The monitoring method based on graph embedding and cross-time segment distribution matching according to the embodiment of the present invention is based on recursive principal component analysis. Monitoring diagram;

[0106] Figure 5 The monitoring method based on graph embedding and cross-time segment distribution matching according to the embodiment of the present invention is based on recursive principal component analysis. Monitoring diagram;

[0107] Figure 6 The monitoring method based on graph embedding and cross-time segment distribution matching in the first embodiment of the present invention is based on graph embedding and cross-time segment distribution matching. Monitoring diagram;

[0108] Figure 7 The monitoring method based on graph embedding and cross-time segment distribution matching in the first embodiment of the present invention is based on graph embedding and cross-time segment distribution matching. Monitoring diagram;

[0109] Figure 8 This is a schematic diagram of the structure of a monitoring device based on graph embedding and cross-time segment distribution matching according to the second embodiment of the present invention;

[0110] Figure 9 A schematic diagram of the structure of a computing terminal provided in Embodiment 3 of the present invention. DETAILED DESCRIPTION

[0111] The present invention will be further described in detail below with reference to the accompanying drawings and examples. It will be understood that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present invention, not all structures.

[0112] Example 1

[0113] Figure 1 This is a flow chart of a monitoring method based on graph embedding and cross-time segment distribution matching according to the first embodiment of the present invention. Figure 1 This monitoring method can monitor industrial processes under non-stationary conditions, thereby improving the fault detection rate. The monitoring method specifically includes the following steps:

[0114] Step 101: Obtain non-stationary industrial process data under normal operating conditions, and obtain time segments with the largest distribution differences by segmenting the data.

[0115] Step 102: Construct a graph embedding cross-time segment distribution matching model; wherein, the graph embedding cross-time segment distribution matching model is constructed by learning projection , the non-stationary industrial process data of different time segments are projected into low-dimensional space, and the distribution matching of different time segments is achieved in the low-dimensional space, while the structure of the non-stationary industrial process data is maintained by constructing a joint objective function.

[0116] Step 103: Construct two global monitoring indicators based on Bayesian inference, and use kernel density estimation to calculate the control limits of the two global monitoring indicators to obtain the control limits of the offline modeling stage; wherein, one global monitoring indicator is used to monitor data changes in the principal component space, and the other global monitoring indicator is used to monitor data changes in the residual space.

[0117] In practical applications, we first collect data from normal operating conditions. To characterize the data's evolving distribution, we segment the data to identify time segments with the greatest distribution differences. We then develop a graph embedding cross-time segment distribution matching model. This model, by constructing a joint objective function, reduces the distribution differences between different time segments while preserving the data structure. Next, we design two global monitoring metrics and calculate their control limits.

[0118] Step 104: New sample data of the industrial process is obtained online, and two global monitoring indicators corresponding to the new sample data are calculated and compared with the control limits in the offline modeling stage to obtain a comparison result, and whether a fault occurs in the industrial process is determined based on the comparison result.

[0119] In actual application, when a new sample is collected, its global monitoring index is calculated and compared with the control limit obtained in the offline modeling stage. If it exceeds the control limit, it indicates that a fault has occurred.

[0120] Preferably, step 101, obtaining industrial process data under normal operating conditions and obtaining time segments with the largest distribution differences by segmenting the data, specifically includes the following steps:

[0121] Step 1011: Obtain non-stationary industrial process data under normal conditions ,in is the dimension of the sample data, is the number of samples, to Indicates 1 to samples.

[0122] Step 1012: normalize the non-stationary industrial process data so that each column of variables has a mean of 0 and a variance of 1. The formula is as follows:

[0123] (1)

[0124] In the above formula express Any sample in is the mean vector of the variables, is the variance vector of the variables.

[0125] Step 1013: Use the objective function to segment the normalized non-stationary industrial process data. The formula is as follows:

[0126] (2)

[0127] In the above formula Used to calculate the number of cumulative summations, Indicates the number of time segments, the subscript in the above formula and Subscripts representing different time segments, and Represents different time segment data, which have significant distribution differences, Represents a time segment The sample points, Represents a time segment The sample points, Represents a time segment The number of samples, Represents a time segment The number of samples, and It is a preset parameter used to avoid unnecessary solutions.

[0128] Step 1014: first, divide the normalized non-stationary industrial process data into 2 to 100 units. share( and ),get There are 1 to 2 split strategies. split points; then, under each segmentation strategy, the distribution difference in different cases is calculated using formula (2), and the segmentation strategy with the largest distribution difference is selected as the final solution.

[0129] In practical applications, after obtaining the time segment with the largest distribution difference, a graph embedding cross-time segment distribution matching model is proposed. , projecting data of different time segments into low-dimensional space, achieving distribution matching of different time segments in low-dimensional space and maintaining the structure of the data.

[0130] Preferably, step 102, constructing a graph embedding cross-time segment distribution matching model, specifically includes the following steps:

[0131] Step 1021: Use the maximum average difference to match the distribution between different time segments, and use the first objective function to reduce the distribution difference of data in different time segments to reduce the non-stationarity of the data. The formula is as follows:

[0132] (3)

[0133] In the above formula and Represents time segments and The sample points, and Represents time segments and Projection; Projection is to project high-dimensional data to low-dimensional data, that is, , It is a high-dimensional sample The optimization problem is written in the form of trace, and the formula is as follows:

[0134] (4)

[0135] (5)

[0136] (6)

[0137] (7)

[0138] In the above formula (4) , Indicates the number of time segments; Indicates the number of elements is The all-one vector of Indicates the number of elements is The all-one vector of 、 、 represents a fully unified matrix, and equations (5), (6), and (7) represent the matrices of any time segment respectively. Seek the Dharma.

[0139] Specifically, for generalization, only the matrices of the first and last time segments are shown. Indicates that ellipsis is used in the middle.

[0140] Step 1022: Construct a second objective function to maintain the local geometric structure of each time segment. The formula is as follows:

[0141] (8)

[0142] In the above formula It's a time segment The adjacency matrix of is as follows:

[0143] (9)

[0144] In the above formula is the thermonuclear parameter, express of The optimization problem is written in the form of traces, and the formula is as follows:

[0145] (10)

[0146] In the above formula Represents the data of the first time segment, Represents the data of the last time segment, represents the Laplacian graph, is a diagonal matrix, represents an all-zero matrix.

[0147] In practical applications, the internal structure of the data may be destroyed during distribution matching, which can affect process monitoring. Therefore, this embodiment uses graph embedding to preserve the geometric structure of the data. The second objective function ensures that samples that are close in high-dimensional space are also close in low-dimensional space, thereby preserving the local geometric structure of the data.

[0148] Step 1023: Construct a third objective function to maintain the global geometric structure of each time segment. The formula is as follows:

[0149] (11)

[0150] In the above formula is the non-adjacency matrix of each time segment, and the formula is as follows:

[0151] (12)

[0152] In the above formula is the thermonuclear parameter, represents the geodesic distance between two sample points in the same time segment; where the optimization problem is expressed in the form of a trace, the formula is as follows:

[0153] (13)

[0154] In the above formula represents the Laplacian graph, is a diagonal matrix.

[0155] In practical applications, the third objective function can make samples that are far apart in high-dimensional space also stay away from each other in low-dimensional space, that is, maintain the global geometric structure of the data.

[0156] Step 1024: Combine the optimization problems corresponding to the first objective function, the second objective function, and the third objective function to construct a graph embedding cross-time segment distribution matching model. The formula is as follows:

[0157] (14)

[0158] In the above formula is the weight coefficient of each item, and the joint objective function is transformed into the following formula:

[0159] (15)

[0160] Step 1025: Solve the joint objective function using Lagrangian technology and convert the optimization problem corresponding to the joint objective function into generalized eigenvalue decomposition to obtain the projection , the formula is as follows:

[0161] (16)

[0162] In the above formula yes The largest eigenvalues, where the number of selected eigenvalues ​​must be less than the sample dimension. In practical applications, yes The projection matrix composed of the eigenvectors corresponding to the largest eigenvalues ​​is obtained The time segments can then be projected into their respective subspaces.

[0163] Preferably, step 103, constructing two global monitoring indicators based on Bayesian inference, and calculating the control limits of the two global monitoring indicators using kernel density estimation to obtain the control limits of the offline modeling stage, includes:

[0164] Step 1031: Use projection Project each time segment into its own subspace to establish local models, and calculate two statistics belonging to the model samples under each local model. The formula is as follows:

[0165] (17)

[0166] (18)

[0167] In the above formula Statistics are used to monitor data changes in the principal element space. Statistics are used to monitor data changes in the residual space. is the low-dimensional representation of the sample, Represents a time segment covariance of the low-dimensional representation, represents the residual space.

[0168] Step 1032: Calculate control limits using kernel density estimation. Represents a time segment of Statistical control limits, and use Represents a time segment of Statistical control limits.

[0169] Step 1033: Bring each sample into The local model calculates statistics and constructs the The global monitoring indicator is composed of the local results, and the probability of each sample belonging to each local model is calculated. The formula is as follows:

[0170] (19)

[0171] In the above formula represents the marginal probability, and are prior probability and conditional probability respectively;

[0172] The prior probability Calculated by the following formula:

[0173] (20)

[0174] Secondly, the conditional probability of the sample under the two statistics is calculated by the following formula:

[0175] (twenty one)

[0176] (twenty two)

[0177] In the above formula express of The statistic conditional probability, express of The statistic conditional probability, express In the time segment model Next Statistics, express In the time segment model Next Statistics, Represents a time segment of Statistical control limits, Represents a time segment of Statistical control limits.

[0178] Step 1034: The two global monitoring indicators are calculated using the following weighted mean, using the following formula:

[0179] (twenty three)

[0180] (twenty four)

[0181] In the above formula express exist Statistics belong to time segments The probability of express exist Statistics belong to time segments probability.

[0182] Step 1035: Use kernel density estimation to calculate the control limits of the two global monitoring indicators to obtain the control limits of the offline modeling stage.

[0183] Preferably, step 104, obtaining new sample data of the industrial process online, calculating two global monitoring indicators corresponding to the new sample data, comparing them with the control limits of the offline modeling stage to obtain a comparison result, and determining whether a fault occurs in the industrial process based on the comparison result, includes:

[0184] Step 1041: Collect new sample data , and normalized.

[0185] Step 1042: Calculate the statistics of the new sample data under each local model. The formula is as follows:

[0186] (25)

[0187] (26)

[0188] In the above formula Is a low-dimensional representation of the new sample data; where .

[0189] Step 1043: Calculate the probability that the new sample data belongs to each local model. The formula is as follows:

[0190] (27)

[0191] Step 1044: Calculate the conditional probability of the new sample data under the two statistics. The formula is as follows:

[0192] (28)

[0193] (29)

[0194] In the above formula express of Statistic conditional probability, express of The conditional probability of a statistic, express In the time segment model Next Statistics, express In the time segment model Next Statistics.

[0195] Step 1045: Calculate the global monitoring index of the new sample data. The formula is as follows:

[0196] (30)

[0197] (31)

[0198] In the above formula express exist Statistics belong to time segments The probability of express exist Statistics belong to time segments probability.

[0199] Step 1046: Determine whether the global monitoring index of the new sample data exceeds the control limit. If it exceeds the control limit, it is fault data.

[0200] Figure 2 This is a schematic diagram of temporal covariate drift in a monitoring method based on graph embedding and cross-time segment distribution matching described in Example 1 of the present invention.

[0201] Figure 3 Schematic diagram of a monitoring framework of a monitoring method based on graph embedding and cross-time segment distribution matching according to the first embodiment of the present invention.

[0202] Exemplarily, a three-phase flow process is used to help understand this embodiment. The three-phase flow process can provide controllable and measurable water, oil and air for the pressurized system. The three-phase flow process operates under varying operating conditions, exhibits obvious non-stationary characteristics, and is a complex industrial equipment. The three-phase flow process consists of pipes with different pore sizes and geometries and a gas-liquid two-phase separator, which can precisely control the flow and separation of fluids. During the experiment, the system has two main process inputs: air and water. By adjusting these two inputs, experimental data can be generated under different operating conditions. Specifically, the system provides a selection of four air flow rates and five water flow rates. These inputs can be combined into 20 different operating conditions, providing a broad data base for research.

[0203] For the three-phase flow process, 10 variables were selected as monitoring variables. In the offline modeling stage, 2860 samples were selected as training data, and the sampling time interval was 8 seconds. In the online monitoring stage, 727 samples were used as test data. The fault occurred in the 196th sample and ended in the 678th sample. The fault changed slightly at the beginning and then became larger and larger. Two methods were implemented in the experiment, namely the method based on recursive principal component analysis and the method proposed in the present invention. The confidence level of the control limits of the two methods was set to 99%.

[0204] See also Figure 2 and Figure 3 ,For the monitoring method proposed in the embodiment of the present invention, the three-phase flow data is divided into 4 time segments in the offline modeling stage, and the number of nearest neighbors is set to 5. All are set to 1, and the data is finally reduced to 5 dimensions. For recursive principal component analysis, the cumulative variance contribution rate is set to 0.8 to determine the number of principal components.

[0205] Figure 4The monitoring method based on graph embedding and cross-time segment distribution matching according to the embodiment of the present invention is based on recursive principal component analysis. Monitoring chart.

[0206] Figure 5 The monitoring method based on graph embedding and cross-time segment distribution matching according to the embodiment of the present invention is based on recursive principal component analysis. Monitoring chart.

[0207] Figure 6 The monitoring method based on graph embedding and cross-time segment distribution matching in the first embodiment of the present invention is based on graph embedding and cross-time segment distribution matching. Monitoring chart.

[0208] Figure 7 The monitoring method based on graph embedding and cross-time segment distribution matching in the first embodiment of the present invention is based on graph embedding and cross-time segment distribution matching. Monitoring chart.

[0209] Figure 4-7 is the monitoring chart of the two methods. The statistic detected a fault starting with the 521st sample, resulting in a fault detection rate of only 28.26%. Recursive principal component analysis only extracts global features of the data, ignoring local information that is useful for process monitoring. Furthermore, as an adaptive method, it may update the model based on fault data, which can affect subsequent process monitoring performance.

[0210] In contrast, the method proposed in this embodiment The indicator exceeded the control limit starting from the 304th sample, and the fault detection rate reached 75.94%. Therefore, the monitoring method proposed in this embodiment can capture data distribution changes, reduce the impact of data distribution changes on process monitoring through distribution matching, and represent local and global information of the data, extracting more comprehensive information.

[0211] The monitoring method based on graph embedding and cross-time segment distribution matching described in this embodiment obtains non-stationary industrial process data under normal operating conditions and obtains time segments with the largest distribution differences by segmenting the data; constructs a graph embedding and cross-time segment distribution matching model; constructs two global monitoring indicators based on Bayesian inference, and uses kernel density estimation to calculate the control limits of the two global monitoring indicators to obtain the control limits of the offline modeling stage; obtains new sample data of the industrial process online, calculates the two global monitoring indicators corresponding to the new sample data, and compares them with the control limits of the offline modeling stage to obtain a comparison result, and determines whether the industrial process has failed based on the comparison result. Compared with the existing technology, it can realize the monitoring of industrial processes under non-stationary conditions and achieve the purpose of improving the fault detection rate.

[0212] Example 2

[0213] Figure 8 This is a structural diagram of a monitoring device based on graph embedding and cross-time segment distribution matching according to the second embodiment of the present invention. Figure 8 , such monitoring device includes:

[0214] The offline module 201 is used to obtain non-stationary industrial process data under normal operating conditions and to obtain time segments with the largest distribution differences by segmenting the data.

[0215] Construction module 202 is used to construct a graph embedding cross-time segment distribution matching model; wherein, the graph embedding cross-time segment distribution matching model is constructed by learning projection , the non-stationary industrial process data of different time segments are projected into low-dimensional space, and the distribution matching of different time segments is achieved in the low-dimensional space, while the structure of the non-stationary industrial process data is maintained by constructing a joint objective function.

[0216] The calculation module 203 is used to construct two global monitoring indicators based on Bayesian inference, and use kernel density estimation to calculate the control limits of the two global monitoring indicators to obtain the control limits of the offline modeling stage; wherein, one global monitoring indicator is used to monitor the data changes in the principal component space, and the other global monitoring indicator is used to monitor the data changes in the residual space.

[0217] The comparison module 204 is used to obtain new sample data of the industrial process online, calculate two global monitoring indicators corresponding to the new sample data, compare them with the control limits of the offline modeling stage to obtain a comparison result, and determine whether the industrial process has a fault based on the comparison result.

[0218] The monitoring device based on graph embedding and cross-time segment distribution matching provided in an embodiment of the present invention can execute the monitoring method based on graph embedding and cross-time segment distribution matching provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0219] Example 3

[0220] Figure 9 A schematic diagram of the structure of a computing terminal provided in the third embodiment of the present invention; Figure 9 A block diagram of an exemplary terminal suitable for implementing embodiments of the present invention is shown. Figure 9 The terminal shown is only an example and should not limit the functions and scope of use of the embodiments of the present invention.

[0221] like Figure 9 As shown, terminal 12 is implemented as a general-purpose computing device. Components of terminal 12 may include, but are not limited to, one or more processors or processing units 16, a system memory 28, and a bus 18 that connects various system components (including system memory 28 and processing unit 16).

[0222] Bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor, or a local bus using any of a variety of bus architectures. Examples of these architectures include, but are not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MAC) bus, an Enhanced ISA bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.

[0223] The terminal 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the terminal 12, including volatile and non-volatile media, removable and non-removable media.

[0224] System memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Terminal 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be configured to read and write non-removable, non-volatile magnetic media ( Figure 9 Not shown, usually called a "hard drive"). Although Figure 9Although not shown, a magnetic disk drive for reading and writing to a removable non-volatile magnetic disk (e.g., a "floppy disk"), as well as an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 18 via one or more data media interfaces. Memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of various embodiments of the present invention.

[0225] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each of which, or some combination thereof, may include an implementation of a network environment. Program modules 42 generally implement the functions and / or methodologies of the embodiments described herein.

[0226] Terminal 12 can also communicate with one or more external devices 14 (e.g., a keyboard, pointing device, display 24, etc.), one or more devices that enable a user to interact with terminal 12, and / or any device that enables terminal 12 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). This communication can occur via input / output (I / O) interface 22. Furthermore, terminal 12 can communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via network adapter 20. As shown, network adapter 20 communicates with other modules of terminal 12 via bus 18. It should be understood that, although not shown, other hardware and / or software modules can be used in conjunction with terminal 12, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0227] The processing unit 16 executes various functional applications and data processing by running programs stored in the system memory 28, such as implementing the monitoring method based on graph embedding and cross-time segment distribution matching provided by an embodiment of the present invention.

[0228] Example 4

[0229] Embodiment 4 of the present invention further provides a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to execute any of the monitoring methods based on graph embedding and cross-time segment distribution matching provided in the above embodiments.

[0230] The computer storage medium of the embodiments of the present invention may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0231] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0232] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0233] Computer program code for performing the operations of the present invention may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0234] Example 5

[0235] The fifth embodiment of the present invention further provides a computer product, including a computer program, which, when executed by a processor, implements the monitoring method based on graph embedding and cross-time segment distribution matching as described in the above embodiment.

[0236] Note that the above are only preferred embodiments of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and that various obvious changes, readjustments, and substitutions can be made by those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments and may include many other equivalent embodiments without departing from the concept of the present invention. The scope of the present invention is determined by the scope of the appended claims.

Claims

1. A monitoring method based on graph embedding and cross-time segment distribution matching, characterized in that: include: Obtain non-stationary industrial process data under normal operating conditions and segment the data to obtain the time segments with the largest distribution differences; Construct a graph embedding cross-time segment distribution matching model; wherein, the graph embedding cross-time segment distribution matching model is constructed by learning projection , the non-stationary industrial process data of different time segments are projected into a low-dimensional space, and the distribution matching of different time segments is achieved in the low-dimensional space, while the structure of the non-stationary industrial process data is maintained by constructing a joint objective function; Two global monitoring indicators are constructed based on Bayesian inference, and the control limits of the two global monitoring indicators are calculated using kernel density estimation to obtain the control limits of the offline modeling stage. Among them, one global monitoring indicator is used to monitor data changes in the principal component space, and the other global monitoring indicator is used to monitor data changes in the residual space. Acquire new sample data of the industrial process online, calculate two global monitoring indicators corresponding to the new sample data, compare them with the control limits in the offline modeling stage to obtain a comparison result, and determine whether the industrial process has a fault based on the comparison result; The construction of the graph embedding cross-time segment distribution matching model includes: The maximum average difference is used to match the distribution between different time segments, and the first objective function is used to reduce the distribution difference of data in different time segments to reduce the non-stationarity of the data. The formula is as follows: (3) In the above formula and Represents time segments and The sample points, and Represents time segments and The optimization problem is expressed in the form of a trace, as follows: (4) (5) (6) (7) In the above formula (4) , Indicates the number of time segments; Indicates the number of elements is The all-one vector of Indicates the number of elements is The all-one vector of 、 、 represents a fully unified matrix, and equations (5), (6), and (7) represent the matrices of any time segment respectively. Seek the Dharma; The second objective function is constructed to maintain the local geometric structure of each time segment. The formula is as follows: (8) In the above formula It's a time segment The adjacency matrix of is as follows: (9) In the above formula is the thermonuclear parameter, express of The optimization problem is written in the form of traces, and the formula is as follows: (10) In the above formula Represents the data of the first time segment, Represents the data of the last time segment, represents the Laplacian graph, is a diagonal matrix, represents an all-zero matrix; The third objective function is constructed to maintain the global geometric structure of each time segment. The formula is as follows: (11) In the above formula is the non-adjacency matrix of each time segment, and the formula is as follows: (12) In the above formula is the thermonuclear parameter, represents the geodesic distance between two sample points in the same time segment; where the optimization problem is expressed in the form of a trace, the formula is as follows: (13) In the above formula represents the Laplacian graph, is a diagonal matrix; Combining the optimization problems corresponding to the first, second, and third objective functions, a graph embedding distribution matching model across time segments is constructed. The formula is as follows: (14) In the above formula is the weight coefficient of each item, and the joint objective function is transformed into the following formula: (15) The joint objective function is solved using the Lagrangian technique, and the optimization problem corresponding to the joint objective function is converted into a generalized eigenvalue decomposition to obtain the projection , the formula is as follows: (16) In the above formula yes The largest eigenvalue, is the corresponding eigenvector, where the number of selected eigenvalues ​​must be less than the sample dimension.

2. The method according to claim 1, characterized in that The method of obtaining industrial process data under normal operating conditions and obtaining time segments with the largest distribution differences by segmenting the data includes: Acquisition of non-stationary industrial process data under normal conditions ,in is the dimension of the sample data, is the number of samples, to Indicates 1 to samples; The non-stationary industrial process data is normalized so that each column of variables has a mean of 0 and a variance of 1. The formula is as follows: (1) In the above formula express Any sample in is the mean vector of the variables, is the variance vector of the variables; The objective function is used to segment the normalized non-stationary industrial process data. The formula is as follows: (2) In the above formula Used to calculate the number of cumulative summations, Indicates the number of time segments, subscript and Subscripts representing different time segments, and Indicates different time segment data, and there are differences in the distribution of the two. Represents a time segment The sample points, Represents a time segment The sample points, Represents a time segment The number of samples, Represents a time segment The number of samples, and It is a preset parameter used to avoid unnecessary solutions; First, the normalized non-stationary industrial process data are divided into 2 to share( and ),get There are 1 to split points; then, under each segmentation strategy, the distribution difference in different cases is calculated using formula (2), and the segmentation strategy with the largest distribution difference is selected as the final solution.

3. The method according to claim 1, characterized in that The two global monitoring indicators are constructed based on Bayesian inference, and the control limits of the two global monitoring indicators are calculated using kernel density estimation to obtain the control limits of the offline modeling stage, including: Using projection Project each time segment into its own subspace to establish local models, and calculate two statistics belonging to the model samples under each local model. The formula is as follows: (17) (18) In the above formula Statistics are used to monitor data changes in the principal element space. Statistics are used to monitor data changes in the residual space. is the low-dimensional representation of the sample, Represents a time segment covariance of the low-dimensional representation, represents the residual space; Calculate the control limits using kernel density estimation, using Represents a time segment of Statistical control limits, and use Represents a time segment of Statistical control limits; Bring each sample into The local model calculates statistics and constructs the The global monitoring indicator is composed of the local results, and the probability of each sample belonging to each local model is calculated. The formula is as follows: (19) In the above formula represents the marginal probability, and are prior probability and conditional probability respectively; The prior probability Calculated by the following formula: (20) Secondly, the conditional probability of the sample under the two statistics is calculated by the following formula: (21) (22) In the above formula express of The statistic conditional probability, express of The statistic conditional probability, express In the time segment model Next Statistics, express In the time segment model Next Statistics, Represents a time segment of Statistical control limits, Represents a time segment of Statistical control limits; The two global monitoring indicators are calculated using the following weighted averages: (23) (24) In the above formula express exist Statistics belong to time segments The probability of express exist Statistics belong to time segments probability; The control limits of the two global monitoring indicators are calculated using kernel density estimation to obtain the control limits in the offline modeling stage.

4. The method according to claim 1, wherein The online acquisition of new sample data of the industrial process, the calculation of two global monitoring indicators corresponding to the new sample data, and the comparison results obtained by comparing the two global monitoring indicators with the control limits of the offline modeling stage are performed, and the determination of whether a fault occurs in the industrial process based on the comparison results includes: Collect new sample data , and normalized; Calculate the statistics of the new sample data under each local model. The formula is as follows: (25) (26) In the above formula Is a low-dimensional representation of the new sample data; where ; Calculate the probability that the new sample data belongs to each local model. The formula is as follows: (27) Calculate the conditional probability of new sample data under two statistics. The formula is as follows: (28) (29) In the above formula express of The statistic conditional probability, express of The conditional probability of a statistic, express In the time segment model Next Statistics, express In the time segment model Next Statistics; Calculate the global monitoring indicators of new sample data using the following formula: (30) (31) In the above formula express exist Statistics belong to time segments The probability of express exist Statistics belong to time segments probability; Determine whether the global monitoring indicators of the new sample data exceed the control limit. If they exceed the control limit, it is fault data.

5. A monitoring device applied to the method according to claim 1, characterized in that: include: The offline module is used to obtain non-stationary industrial process data under normal operating conditions and to obtain the time segments with the largest distribution differences by segmenting the data; A construction module is used to construct a graph embedding cross-time segment distribution matching model; wherein, the graph embedding cross-time segment distribution matching model is constructed by learning projection , the non-stationary industrial process data of different time segments are projected into a low-dimensional space, and the distribution matching of different time segments is achieved in the low-dimensional space, while the structure of the non-stationary industrial process data is maintained by constructing a joint objective function; The calculation module is used to construct two global monitoring indicators based on Bayesian inference and calculate the control limits of the two global monitoring indicators using kernel density estimation to obtain the control limits in the offline modeling stage. Among them, one global monitoring indicator is used to monitor data changes in the principal component space, and the other global monitoring indicator is used to monitor data changes in the residual space. The comparison module is used to obtain new sample data of the industrial process online, calculate the two global monitoring indicators corresponding to the new sample data, compare them with the control limits of the offline modeling stage to obtain comparison results, and judge whether the industrial process has failed based on the comparison results.

6. A terminal, characterized in that: include: one or more processors; a storage device for storing one or more programs; Sensors, used to collect data; When the one or more programs are executed by the one or more processors, the one or more processors implement the monitoring method based on graph embedding and cross-time segment distribution matching as described in any one of claims 1 to 4.

7. A storage medium containing computer-executable instructions, characterized in that: When the computer executable instructions are executed by a computer processor, the computer executable instructions are used to perform the monitoring method based on graph embedding and cross-time segment distribution matching according to any one of claims 1 to 4.

8. A computer product, comprising a computer program, characterized in that: When the computer program is executed by a processor, the monitoring method based on graph embedding and cross-time segment distribution matching as described in any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Fault detection method based on double-layer distributed and integrated Bayesian reasoning

    CN117093932A

  • Intrusion detection system and method based on intelligent network

    CN118413406A