Uncertain process fault monitoring method based on local and global interval embedding
By constructing local and global interval embedding models and combining complete information principal component analysis and interval local preservation projection, the problem of insufficient utilization of local neighborhood information in traditional methods is solved, achieving more efficient fault monitoring and accurate location, and improving the accuracy and sensitivity of fault detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TAIYUAN UNIVERSITY OF TECHNOLOGY
- Filing Date
- 2023-08-15
- Publication Date
- 2026-05-29
AI Technical Summary
Traditional interval principal component analysis methods fail to effectively utilize local neighborhood information when dealing with measurement data affected by uncertain factors, resulting in insufficient accuracy and sensitivity in fault monitoring and difficulty in accurately locating fault variables.
By combining complete information principal component analysis with interval local preservation projection, local and global interval embedding models are constructed. By minimizing the projection of local scattering and maximizing the projection of global scattering, feature information of interval value data is extracted, and kernel density estimation is used to determine control limits and monitor faults in data samples.
It improves the accuracy and sensitivity of fault monitoring, can accurately locate fault variables, has a good fault isolation effect, reduces false alarm rate and false negative rate, and improves the accuracy and generalization ability of data classification.
Smart Images

Figure CN117032170B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial process fault monitoring, and specifically to a method for monitoring uncertain process faults based on local and global interval embedding. Background Technology
[0002] With the ever-increasing demand for process safety and high-quality products, process fault monitoring has become an essential part of the daily operation of many industrial processes. In fact, some process faults can disrupt the entire system's production, leading to loss of property and life. Therefore, in order to monitor equipment faults in a timely and accurate manner, an efficient fault detection and isolation algorithm is of great significance for achieving high safety and reliability in many process control systems.
[0003] With the widespread application of data acquisition technology, various industrial processes have recorded a large amount of historical data. Methods based on multivariate statistical process control have developed rapidly; however, the reliability of traditional methods largely depends on the quality of the measurement data. In actual industrial processes, measurement data is often affected by various uncertainties. To address these issues, interval representation has become an effective method for handling such uncertainties.
[0004] Traditional interval principal component analysis (PCA) only considers global information of interval-valued data, rarely taking into account local neighborhood information. However, local neighborhood information can characterize the topological relationships between interval-valued data points and uncover meaningful low-dimensional information hidden within high-dimensional process data. Therefore, this paper combines full-information PCA with interval local-preserving projection to extract more useful features from interval-valued data. Summary of the Invention
[0005] In order to simultaneously preserve local neighborhood and global information of interval value data, improve the accuracy and sensitivity of interval value data fault monitoring, accurately locate fault variables, and achieve good fault isolation effect, this invention provides a local and global interval embedding algorithm applicable to uncertain system processes.
[0006] This invention adopts the following technical solution: a fault monitoring method for uncertain processes based on local and global interval embedding, comprising:
[0007] S1: Obtain an inaccurate single-value dataset of the device under normal conditions. The dataset includes a combination of actual collected data and noise uncertainty factors. The dataset is then preprocessed to obtain a standardized interval value dataset.
[0008] S2: Establish local and global interval embedding models and apply them to feature extraction of interval value data to determine interval principal components;
[0009] S3: Calculate the statistics associated with the normal interval value dataset based on the principal components of the intervals, and use kernel density estimation to determine the control limits of the statistics;
[0010] S4: Collect new data samples, and preprocess the newly collected data samples to obtain standardized interval values of the new data samples.
[0011] S5: Calculate the statistics of the new interval value data using the interval principal components determined by the local and global interval embedding models;
[0012] S6: Monitor whether the four newly obtained statistics exceed the control limits. If they do, the system malfunctions and proceeds to step S7. Otherwise, return to step S4 and monitor the next sample.
[0013] S7: If a fault occurs at a new sample, the variable that contributes highly to the fault in the contribution plot is the fault variable.
[0014] In some embodiments, step S1 includes:
[0015] S11: Calculate the relative measurement error by using the deviation between the actual value and the measured value of the variable. The variable is the value of the variable in the sample from the collected imprecise single-valued dataset;
[0016] S12: Determine the lower and upper limits of the relative measurement error. and And it transforms the uncertain data it measures into interval value data;
[0017] S13: Standardize the interval value data.
[0018] In some embodiments, in step S11, if the actual value of the j-th variable for all samples cannot be measured, the upper and lower limits are determined based on expert experience or the measurement error provided by the sensor manufacturer.
[0019] In some embodiments, in step S12: if the actual value of the j-th variable for all samples can be measured in a laboratory or similar manner, then the relative measurement is calculated. ,in and Let represent the uncertain measured value and the actual value of the j-th process variable, respectively. Then, the most reasonable lower and upper bounds of the relative measurement error of the j-th process variable are determined by a measurement error estimation method based on the principle of reasonable granularity. The uncertain data of the i-th observation of the j-th process variable is obtained through... and Convert to range data .
[0020] In some embodiments, the measurement error estimation method based on the principle of reasonable granularity obtains a reasonable range for uncertain data by balancing the conflicting requirements of coverage and specificity, and maximizing their product. The formula for determining the reasonable range is as follows: ,in, Indicates coverage rate. Indicates specificity, 'a' represents the lower bound of the interval, and 'b' represents the upper bound of the interval.
[0021] In some embodiments, step S2 includes:
[0022] S21: Constructing an interval value dataset Complete information principal component analysis covariance matrix :
[0023] S22: Calculate the local similarity matrix W Then calculate the diagonal matrix. and Laplace matrix ;
[0024] S23: Construct an interval value dataset using the parameters from step S22. Local matrix of local preservation projection in interval U ;
[0025] S24: Establish the objective function by minimizing local scattering and maximizing global scattering. Minimize the objective function to calculate the eigenvalues and the corresponding feature vector ;
[0026] S25: Determine the prior based on the cumulative percentage variance standard. l principal component vectors And calculate the input interval value data. Projection features Estimated value and interval residual matrix .
[0027] In some embodiments, in step S21 The calculation process is as follows:
[0028] when hour,
[0029] ;
[0030] when hour,
[0031] ;
[0032] in, and This represents a dataset of interval values. The lower and upper bounds of the interval values in the k-th row and i-th column. Indicates the inner product. Representing a range value dataset The value of the i-th variable.
[0033] In some embodiments, in step S22, the local similarity matrix W Calculated using the following formula:
[0034] ;
[0035] in and Let these be the lower and upper bounds of the i-th sample in the interval-valued data. It is an empirical constant. Represents the interval value sample It consists of k samples with the minimum Euclidean distance, where , This is the adjacency parameter.
[0036] In some embodiments, the local matrix in step S23 U for,
[0037] Given two interval-valued variables and ,when When, it is defined as:
[0038]
[0039] when When, it is defined as:
[0040] .
[0041] In step S4, there are four statistics, namely: , , as well as Kernel density estimation was used to determine the control limits for the four statistics. , , and .
[0042] Compared with the prior art, the present invention has the following beneficial effects:
[0043] (1) By using the measurement error estimation method based on the principle of reasonable granularity, it is possible not only to effectively convert inaccurate data with high measurement noise and large measurement error into interval value data, but also to fully reflect the data characteristics, eliminate the interference of outliers, and obtain the optimal interval of inaccurate data.
[0044] (2) A feature extraction method employing local and global interval embedding algorithms finds a projection that simultaneously preserves local and global information by minimizing local scattering and maximizing global scattering, effectively extracting feature information from interval-valued data. Compared to traditional interval principal component analysis, this method can capture more meaningful local neighborhood information while preserving global information from interval-valued data. Therefore, it is more powerful than traditional interval principal component analysis in extracting useful information from interval-valued data, demonstrating more reliable and robust fault detection performance.
[0045] (3) By defining four monitoring and statistical indicators and providing corresponding fault diagrams, it is possible to analyze the running status of the process more comprehensively and accurately identify fault variables, thus achieving a good fault isolation effect. Attached Figure Description
[0046] Figure 1 This is a flowchart of the present invention. Detailed Implementation
[0047] To make the above features and advantages of the present invention more apparent and understandable, the present invention will be clearly and in detail described below with reference to the accompanying drawings and embodiments. This embodiment is based on the technical solution of the present invention and provides a specific implementation process and monitoring scheme, but the present invention is not limited to the specific embodiments disclosed below.
[0048] Example 1: As Figure 1 As shown, the uncertain process fault monitoring method based on local and global interval embedding of the present invention includes the following steps:
[0049] S100: Obtain the inaccurate single-value dataset of the device under normal conditions, and preprocess the dataset to obtain a standardized interval value dataset.
[0050] In this embodiment, the local and global interval embedding algorithms are validated using the Tennessee-Eastman process simulation dataset.
[0051] The Tennessee-Eastman process dataset contains 21 failure modes, presenting different types of process failures. Each failure dataset contains 960 samples, with process disturbances starting from the 161st sample. In this study, data samples under normal operating conditions were used first, and all values obtained from these samples were considered accurate. The actual data is assumed to be... Inaccurate measurement data is represented as The measurement error is Obtaining inaccurate measurement data Measurement error ,in , It follows a Gaussian distribution with a mean of 0 and a standard deviation of 0.01.
[0052] Calculate the relative measurement error by considering the deviation between the actual and measured values of the variable. Represented as in, It is the first j Uncertain measured values of process variables, It is the first j The actual values measured for each process variable.
[0053] The lower and upper limits of the relative measurement error are determined using a measurement error estimation method based on the principle of reasonable granularity. and The uncertain data of the i-th observation of the j-th process variable can be obtained through... and Convert to range data .
[0054] For the above interval value data Standardization is performed, and the standardization formula is as follows:
[0055]
[0056] in and These are the training set numbers. The mean and variance of an interval variable, for an interval variable Their mean and variance are respectively and .
[0057] S200: Establish local and global interval embedding models and apply them to feature extraction of interval value data to determine interval principal components.
[0058] The local and global interval embedding algorithm is driven by a popular learning algorithm that combines complete information principal component analysis and interval local-preserving projection. By minimizing local scattering while maximizing global scattering, it projects high-dimensional interval data into a low-dimensional space, revealing the inherent geometric structure of imprecise data and obtaining more useful low-dimensional information hidden in high-dimensional data. The specific process is as follows:
[0059] First, construct the interval value training dataset. Complete information principal component analysis covariance matrix as follows:
[0060]
[0061] Given two interval-valued variables and ,when When, it is defined as:
[0062]
[0063] when When, it is defined as:
[0064] .
[0065] Next, calculate the local similarity matrix. W The definition is as follows:
[0066]
[0067] Find the diagonal matrix and Laplace matrix .
[0068] Construct an interval value dataset using the parameters described above. Local matrix of local preservation projection in interval U The calculation is as follows:
[0069]
[0070] Given two interval-valued variables and ,when When, it is defined as:
[0071]
[0072] when When, it is defined as:
[0073]
[0074] Again, by minimizing local scattering And maximize global scattering The objective function is established by the projection, and its formula is as follows.
[0075] .
[0076] The above minimization problem can be transformed into the following generalized eigenvalue problem.
[0077]
[0078] Finally, calculate the eigenvalues and eigenvectors of the generalized eigenvalue problem, and let the eigenvectors be... Based on eigenvalues Sort the components from largest to smallest and determine the principal components using the cumulative percentage variance standard.
[0079] Projected score matrix of interval data The calculation is as follows:
[0080] ,
[0081] Estimated value of initial interval data The calculation is as follows:
[0082] ,
[0083] Where the matrix , and These are the initial data matrices. The lower and upper bounds of the estimated value.
[0084] S300: Calculate four statistics associated with the normal interval value dataset based on the interval principal components, and determine the control limits of the four statistics using kernel density estimation.
[0085] The four statistics are two each. and two , A statistic is a measure of the sum of the spatial variation of the latent variable and the normalized fractional squares. The indicator is given by the following formula.
[0086] ,
[0087] in and Representing the projection score matrix respectively The upper and lower limits, It is the normal operating condition projection score matrix The inverse of the covariance matrix, i.e. , .
[0088] The index, also known as the squared prediction error, represents the Euclidean distance between the measured and estimated values in an interval, expressed through the sample vector in the interval residual matrix. Projection measurement on the surface. For Residual features of the i-th interval value dataset ,two The indicator is given by the following formula.
[0089] ,
[0090] Determined by the following kernel density estimation method , , and The control limit is , , and .
[0091] , , , ,in, The upper quantile is the significance level.
[0092] S400: Collect new data samples and preprocess them to obtain standardized interval value data.
[0093] The fault data in the Tennessee-Eastman process are preprocessed using the method in step S100 to obtain standardized interval value data. .
[0094] S500: Calculates four statistics for new interval-valued data using the interval principal components determined by the local and global interval embedding models.
[0095] Apply the interval principal components determined by the model to the interval value data. Four statistical measures , , or It is given by the following formula.
[0096]
[0097]
[0098] ,
[0099] in, and These represent the new interval values. The lower and upper bounds of the i-th variable, and As given in step S300, Let be the i-th value of the j-th eigenvector. and These represent the lower and upper bounds of the projection matrix for interval-valued data, respectively. , , and These represent the i-th variable pairs of the interval value data. , , and The contribution value.
[0100] Monitor whether the four newly obtained statistics exceed the control limits. If they do, the system malfunctions and proceeds to step S700; otherwise, proceed to step S800.
[0101] monitor , , or Does it exceed the control threshold? , , or If all four control indicators are within the threshold range, the system is in normal condition; otherwise, the system malfunctions.
[0102] This paper compares the proposed method with traditional interval principal component analysis (PCA) using three model evaluation metrics: false alarm rate (FAR), false negative rate (MAR), and accuracy (ACC), as shown in Table 1. The table shows that the proposed local neighborhood and global interval PCA methods have relatively low false alarm and false negative rates, and the highest accuracy, with an average accuracy of 93.65%.
[0103] Table 1. Summary of FAR, MDR, and ACC values (%) for the TEP dataset
[0104]
[0105] The S700 calculates a contribution map to identify fault variables.
[0106] When a system malfunctions, it is necessary to analyze the contribution of each process variable to the malfunction statistics. Therefore, contribution plots are used for fault isolation in industrial processes. In other words, a high contribution of a variable to the statistical data means that there is a problem with that variable. , , or The contributions of the statistics are respectively from , , and Perform the calculation.
[0107] Return to step S400 and monitor the next sample.
[0108] If the sample is in a normal state, then monitor the next sample.
[0109] Example 2: To further verify the applicability of the model, we used the coal mining machine operation dataset provided by the Xiegou Coal Mine of Shanxi Xishan Jinxing Energy Co., Ltd. This dataset consists of 1500 normal data samples collected from 11:30 AM to 12:00 PM on March 4, 2019, averaging 50 data samples per minute. It includes 20 monitored variables such as coal mining machine voltage, current, and temperature. The 10 collected samples and the 20 variables they contain are shown in Table 2. 500 samples were used as training samples, and 1000 as test samples. Due to the harsh working environment of the coal mining machine, the collected data is often affected by various factors, greatly impacting data quality. Therefore, this paper uses interval value data to represent uncertain data and uses a measurement error estimation method based on the principle of reasonable granularity to transform inaccurate measurement data into interval value data. The false alarm rate, missed detection rate, and accuracy of coal mining machine fault monitoring obtained through MRPCA, CIPCA, and LG-IPCA models are shown in Table 3.
[0110] Table 2 shows the 10 samples of coal mining machines collected and the 20 variables included therein.
[0111]
[0112]
[0113] Table 3 Summary of FAR, MDR, and ACC values for coal mining machine data (%)
[0114]
[0115] As shown in the table, although the false alarm rates of the three models are almost the same, the MRPCA and CIPCA methods failed to detect the fault after it occurred, with extremely high false alarm rates of 60.25% and 45%, respectively. In contrast, the LG-IPCA method quickly exceeds the threshold after the fault occurs, with a false alarm rate of only 0.25%, and its fault detection accuracy is significantly improved, reaching 98.7%. Therefore, the LG-IPCA method can effectively extract basic process features covered by inaccurate data, greatly improving the monitoring performance of the coal mining machine's process status.
[0116] This invention employs a measurement error estimation method based on a reasonable granularity principle. This method not only effectively converts imprecise data with high measurement noise and large measurement errors into interval-valued data, but also fully reflects data characteristics, eliminates outlier interference, and obtains the optimal interval for imprecise data. The local neighborhood and global interval principal component analysis method finds a projection that simultaneously preserves local and global information by minimizing local scattering and maximizing global scattering. It retains the global structure of interval-valued data while capturing more meaningful local neighborhood information, effectively extracting feature information from the interval-valued data. The defined four monitoring statistical indicators can more comprehensively analyze the process's operational status, and the contribution graph can effectively identify fault variables. This not only significantly reduces false alarm and false negative rates and improves fault detection accuracy and sensitivity, but also accurately locates fault variables, exhibits good fault isolation effects, improves data classification accuracy, and demonstrates good generalization ability.
Claims
1. A method for monitoring faults in uncertain processes based on local and global interval embedding, characterized in that, include: S1: Obtain an inaccurate single-value dataset of the device under normal conditions. The dataset includes a combination of actual collected data and noise uncertainty factors. The dataset is then preprocessed to obtain a standardized interval value dataset. S2: Establish local and global interval embedding models and apply them to feature extraction of interval value data to determine interval principal components; S3: Calculate the statistics associated with the normal interval value dataset based on the principal components of the intervals, and use kernel density estimation to determine the control limits of the statistics; S4: Collect new data samples, and preprocess the newly collected data samples to obtain standardized interval values of the new data samples. S5: Calculate the statistics of the new interval value data using the interval principal components determined by the local and global interval embedding models; S6: Monitor whether the four newly obtained statistics exceed the control limits. If they do, the system malfunctions and proceeds to step S7. Otherwise, return to step S4 and monitor the next sample. The statistics are respectively , , as well as Kernel density estimation was used to determine the control limits for the four statistics. , , and ; S7: If a fault occurs at a new sample, the variable that contributes highly to the fault in the contribution plot is the fault variable.
2. The method for monitoring uncertain process faults based on local and global interval embedding according to claim 1, characterized in that, Step S1 includes: S11: Calculate the relative measurement error by the deviation between the actual value and the measured value of the variable. The variable is the value of the variable in the sample from the collected imprecise single-valued dataset; S12: Determine the lower and upper limits of the relative measurement error. and And it transforms the uncertain data it measures into interval value data; S13: Standardize the interval value data.
3. The method for monitoring uncertain process faults based on local and global interval embedding according to claim 2, characterized in that, In step S11, if the actual value of the j-th variable for all samples cannot be measured, the upper and lower limits are determined based on expert experience or the measurement error provided by the sensor manufacturer.
4. The method for monitoring uncertain process faults based on local and global interval embedding according to claim 2, characterized in that, In step S12: if the actual value of the j-th variable in all samples can be measured in a laboratory manner, then calculate the relative measurement error. ,in and Let represent the uncertain measured value and the actual value of the j-th process variable, respectively. Then, the most reasonable lower and upper bounds of the relative measurement error of the j-th process variable are determined by a measurement error estimation method based on the principle of reasonable granularity. The uncertain data of the i-th observation of the j-th process variable is obtained through... and Convert to range data .
5. The method for monitoring uncertain process faults based on local and global interval embedding according to claim 4, characterized in that, The measurement error estimation method based on the principle of reasonable granularity balances the conflicting requirements of coverage and specificity, maximizing their product to obtain a reasonable interval for uncertain data. The formula for determining the reasonable interval is as follows: ,in, Indicates coverage rate. Indicates specificity, 'a' represents the lower bound of the interval, and 'b' represents the upper bound of the interval.
6. The method for monitoring uncertain process faults based on local and global interval embedding according to claim 1, characterized in that, Step S2 includes: S21: Constructing an interval value dataset Complete information principal component analysis covariance matrix : S22: Calculate the local similarity matrix W Then calculate the diagonal matrix. and Laplace matrix ; S23: Construct an interval value dataset using the parameters from step S22. Local matrix of local preservation projection in interval U ; S24: Establish the objective function by minimizing local scattering and maximizing global scattering. Minimize the objective function to calculate the eigenvalues and the corresponding feature vector ; S25: Determine the preceding value based on the cumulative percentage variance standard. l principal component vectors And calculate the input interval value data. Projection features Estimated value and interval residual matrix .
7. The method for monitoring uncertain process faults based on local and global interval embedding according to claim 6, characterized in that, In step S21 The calculation process is as follows: when hour, ; when hour, ; in, and Representing a range value dataset The lower and upper bounds of the interval values in the k-th row and i-th column. Indicates the inner product. Representing a range value dataset The value of the i-th variable.
8. The method for monitoring uncertain process faults based on local and global interval embedding according to claim 6, characterized in that, In step S22, the local similarity matrix W Calculated using the following formula: ; in and Let these be the lower and upper bounds of the i-th sample in the interval-valued data. It is an empirical constant. Represents the interval value sample It consists of k samples with the minimum Euclidean distance, where , This is the adjacency parameter.
9. The method for monitoring uncertain process faults based on local and global interval embedding according to claim 6, characterized in that, Local matrix in step S23 U for, Given two interval-valued variables and ,when When, it is defined as: when When, it is defined as: 。