A fault detection method and system based on correlation score matrix and support vector data description

CN122839092APending Publication Date: 2026-09-29JIANGNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610977356.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-02
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

这种基于累计方差贡献率的筛选方式存在明显缺陷:特征值的大小与得分向量对故障的敏感程度并不呈线性对应关系

Benefits of technology

本发明的有益效果是:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122839092A_ABST
    Figure CN122839092A_ABST
Patent Text Reader

Abstract

This invention discloses a fault detection method and system based on correlation score matrix and support vector data description. The method includes: performing principal component analysis on the acquired normal operation process data to obtain a full-dimensional score matrix; introducing the concept of variable coupling strength, calculating the correlation between each score vector and the original process variables, and selecting sensitive score vectors based on the correlation to construct a correlation score matrix; using the extracted correlation score matrix as a feature input to the support vector data description model, constructing a minimum hypersphere boundary in the feature space to enclose normal samples as a control limit for fault discrimination; and calculating the distance from the correlation score vector of the test sample to the center of the hypersphere online to achieve fault detection. This invention effectively overcomes the shortcomings of traditional principal component analysis, which relies solely on variance contribution rate to select principal components, easily retaining redundant information and ignoring fault-sensitive information. It also adapts to the nonlinearity and non-Gaussianity of industrial processes, significantly improving the fault detection rate and reducing the false alarm rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial process monitoring and fault detection technology, specifically to a fault detection method and system based on correlation score matrix and support vector data description. Background Technology

[0002] With the rapid development of modern industrial production towards large-scale, continuous, highly integrated, and intelligent operations, advanced distributed control systems are widely used, resulting in the generation of massive amounts of high-dimensional process data during industrial system operation. Accurately extracting effective feature information from this data to achieve real-time monitoring of the production process status and early warning of faults is a core issue in ensuring the safe and stable operation of industrial systems and reducing production losses.

[0003] Principal Component Analysis (PCA), a classic data-driven fault detection method, maps high-dimensional data to a low-dimensional subspace through linear transformation, achieving dimensionality reduction while preserving most of the data's information, and is widely used in industrial processes. However, traditional PCA methods typically select the top components based on a pre-defined cumulative variance contribution rate (such as 85% or 90%). The eigenvectors corresponding to the largest eigenvalues ​​form the principal component loading matrix, which yields the score matrix. This screening method based on cumulative variance contribution rate has a significant drawback: the magnitude of the eigenvalues ​​does not linearly correlate with the sensitivity of the score vectors to faults. When a fault occurs, score vectors with larger eigenvalues ​​may not be sensitive to fault changes, while score vectors with smaller eigenvalues ​​may fluctuate significantly with the fault. Retaining score vectors solely based on variance contribution rate easily retains redundant information unrelated to the fault, while ignoring score vectors that truly contain critical fault information, leading to severe underreporting of faults.

[0004] To address the difficulties faced by centralized monitoring methods in processing high-dimensional and complex data, decentralized industrial process monitoring methods have been extensively studied. The basic idea of ​​these methods is to first decompose process variables into multiple low-dimensional variable sub-blocks according to certain correlation criteria, and then build a monitoring model for each sub-block to reduce monitoring complexity. For example, by constructing a graph model based on mutual information to partition the variable set, and independently training a Support Vector Data Description (SVDD) model on each subset, it can adapt to the non-Gaussian and nonlinear characteristics of the data to a certain extent. However, decentralized methods have the following shortcomings: First, the effectiveness of variable decomposition is highly dependent on the rationality of prior process knowledge or data-driven criteria. If the sub-block partitioning is inappropriate, the coupling relationship between key variables will be severed, making it difficult to effectively capture fault features across sub-blocks. Second, these methods essentially perform subset selection within the original variable space, without quantitatively evaluating and screening the sensitivity of features to faults. Therefore, redundant information unrelated to faults may still be introduced, affecting the accuracy and robustness of fault detection.

[0005] Furthermore, actual industrial process data is affected by external disturbances and operating condition switching, generally exhibiting nonlinear coupling and non-Gaussian distribution characteristics. Support Vector Data Description (SVDD), as a single-classification model, does not have strict prior assumptions about data distribution. By constructing a minimum hypersphere in the feature space through a kernel function, it can effectively handle non-Gaussian and nonlinear problems. However, directly inputting high-dimensional raw data or unfiltered redundant features into the SVDD model not only significantly increases computational complexity, but the redundant information also severely interferes with the construction of decision boundaries, causing the model to fail to accurately learn the boundaries of normal operating conditions.

[0006] Therefore, how to accurately extract the core features that are truly sensitive to faults during the dimensionality reduction process, avoid interference from redundant information, and combine nonlinear single-classification models to improve the fault detection performance of complex industrial processes is an urgent problem to be solved. Summary of the Invention

[0007] To address the shortcomings of existing traditional principal component analysis (PCA) methods, which select principal components solely based on variance contribution rates during feature extraction, potentially retaining redundant information irrelevant to faults while neglecting truly fault-sensitive score vectors, and existing distributed monitoring methods that rely on prior knowledge or data-driven criteria for sub-block division, possibly severing cross-block variable coupling relationships and failing to quantitatively evaluate the fault sensitivity of the features themselves, and to adapt to the nonlinear and non-Gaussian distribution characteristics prevalent in industrial processes, this invention proposes a fault detection method and system based on correlation score matrix and support vector data description (CSM-SVDD). This method calculates the correlation between score vectors and original process variables, uses this correlation to select fault-sensitive score vectors to construct a correlation score matrix, and inputs this matrix as a feature into the support vector data description model for decision boundary learning. This achieves accurate extraction of fault-sensitive features and a significant improvement in fault detection rate.

[0008] To achieve the above objectives, the technical solution adopted by the present invention is as follows: In a first aspect, the present invention provides a fault detection method based on a correlation score matrix and support vector data description, comprising the following steps: Step 1: Obtain historical data samples of the industrial process under normal operating conditions, construct the original dataset and standardize it to obtain a standard dataset with a mean of 0 and a variance of 1, and use the standard dataset as the training set; Step 2: Calculate the covariance matrix corresponding to the standard dataset, and perform eigenvalue decomposition on the covariance matrix to obtain all eigenvalues ​​and the eigenvectors corresponding to each eigenvalue. The projection matrix is ​​constructed from all eigenvectors. The standard dataset is mapped through the projection matrix to obtain the full-dimensional score matrix. Step 3: Calculate the correlation between each score vector in the full-dimensional score matrix and each original process variable of the industrial process, and construct a correlation matrix between variables and score vectors; for each score vector, calculate the sum of the absolute values ​​of the correlation between the score vector and all process variables, and sort them in descending order according to the magnitude of the sum of absolute values; select the top-ranked score vectors as fault-sensitive score vectors according to a preset selection number, combine them to construct a correlation score matrix, and extract the feature vectors corresponding to the fault-sensitive score vectors from the projection matrix to construct a correlation projection matrix; Step 4: Use the relevant score matrix as the model input feature to construct a support vector data description model. Map the low-dimensional features to a high-dimensional space through nonlinear mapping. Solve for the minimum hypersphere that can wrap the normal training samples. Determine the center and radius of the minimum hypersphere. Set the radius of the hypersphere as the control limit for fault detection. Step 5: In the online detection phase, the current test data samples of the industrial system are collected in real time. The mean and variance parameters of the standardization process in Step 1 are used to standardize the test data samples. The relevant projection matrix in Step 3 is called to calculate the relevant score vector corresponding to the standardized test data samples. Step Six: Input the relevant score vector corresponding to the data sample to be tested into the trained support vector data description model, calculate the distance from the center of the smallest hypersphere after the relevant score vector is mapped to the high-dimensional feature space, and use this distance as the fault detection statistic; compare the fault detection statistic with the control limit. If the fault detection statistic is greater than the control limit, it is determined that the industrial system has failed; otherwise, it is determined to be operating normally.

[0009] In one embodiment of the present invention, the standardization process described in step one is performed according to the following formula:

[0010] in, The original value of the variable. The mean of the variable. The standard deviation is the variable.

[0011] In one embodiment of the present invention, the specific process of calculating the covariance matrix, decomposing the eigenvalues, and solving the full-dimensional score matrix in step two is as follows: Let the standard dataset be Calculate the covariance matrix , n The number of training set samples; for the covariance matrix Perform eigenvalue decomposition ,in It is a diagonal matrix composed of eigenvalues ​​sorted in descending order. The projection matrix is ​​composed of the corresponding feature vectors; the formula for calculating the full-dimensional score matrix is: , This is the full-dimensional score matrix obtained after mapping.

[0012] In one embodiment of the present invention, the specific process of calculating the relevance, filtering the fault-sensitive score vectors, and constructing the relevance score matrix in step three is as follows: The Pearson correlation coefficient was used to calculate the first... Original process vectors With the score vectors Relevance A two-dimensional correlation matrix is ​​formed. ; Regarding the first For each score vector, calculate the sum of the absolute values ​​of its correlation with all process variables. ,in v The total number of industrial process variables; by The score vectors are sorted in descending order of value. A predetermined number of score vectors are selected, and the top-ranked vectors are used as fault-sensitive score vectors. These vectors are then combined to construct a relevant score matrix. Extract the feature vectors corresponding to the fault-sensitive score vectors from the projection matrix, and combine them to construct the relevant projection matrix. .

[0013] In one embodiment of the present invention, the specific process of constructing and solving the support vector data description model in step four is as follows: With the relevant score matrix As training samples, a support vector data description model is constructed to optimize the objective, which aims to minimize the volume of the hypersphere. Normal samples are constrained to fall inside the hypersphere or on its surface. The optimization expression is as follows:

[0014] in, The radius of the hypersphere The center of the hypersphere, The penalty coefficient is... As slack variables, The number of training samples. It is a nonlinear mapping function. For the th in the relevant score matrix One sample; By introducing Lagrange multipliers, the constrained optimization problem is transformed into a Lagrange dual optimization problem for solution. A Gaussian kernel function is chosen to replace the high-dimensional inner product operation; the expression for the Gaussian kernel function is: ,in Let the Gaussian kernel width parameter be denoted; solve the dual optimization problem to obtain the center of the hypersphere. With the radius of the hypersphere , radius As a fault detection and control limit.

[0015] In one embodiment of the present invention, the formula for calculating the relevant score vector of the data sample to be tested in step five is as follows:

[0016] in, For the standardized test data, This represents the relevant score corresponding to the data to be tested.

[0017] In one embodiment of the present invention, the logic for calculating fault detection statistics and determining faults in step six is ​​as follows: The formula for calculating the fault detection statistic of the sample to be tested is: ; The logic for determining whether a fault has occurred is as follows: if the fault detection statistic of the sample under test... If the condition is met, the current industrial system is determined to be faulty; otherwise, the system is determined to be in normal operating condition.

[0018] Secondly, the present invention provides a fault detection system based on correlation score matrix and support vector data description, for implementing the aforementioned fault detection method based on correlation score matrix and support vector data description, the system comprising: The data acquisition and standardization preprocessing module is used to collect normal historical data and online test data of industrial processes, and to complete the standardization process according to preset rules to generate training sets and standardized test datasets respectively. The covariance eigenvalue decomposition and score matrix solving module is used to calculate the covariance matrix of the standard dataset, perform eigenvalue decomposition on the covariance matrix to obtain the projection matrix, and combine the projection matrix mapping to solve the full-dimensional score matrix. The sensitive feature filtering module is used to calculate the correlation between the score vector and the original process variable, sort them in descending order according to the sum of the absolute values ​​of the correlation, filter the fault-sensitive score vectors according to the preset selection quantity, and construct the correlation score matrix and the correlation projection matrix. The SVDD offline modeling module is used to train a support vector data description model with the relevant score matrix as input, solve for the center and radius of the minimum hypersphere, and determine the fault detection control limits. The online feature calculation module is used to call the relevant projection matrix and calculate the relevant score vector corresponding to the standardized test data. The fault detection module is used to calculate the fault detection statistics of the sample under test, compare the statistics with the control limits, and output the operating status judgment result of the industrial system.

[0019] Thirdly, the present invention provides a computer-readable storage medium storing computer instructions, which are executed by a processor to describe the fault detection method based on correlation score matrix and support vector data.

[0020] Fourthly, the present invention provides a computer program product storing computer instructions, which are executed by a processor using the fault detection method described above based on correlation score matrix and support vector data. The beneficial effects of this invention are: 1. This invention proposes a fault detection method based on correlation score matrices and support vector data description. It addresses the shortcomings of traditional principal component analysis (PCA), which selects principal components solely based on eigenvalue magnitude, easily retaining redundant information unrelated to the fault while ignoring truly fault-sensitive score vectors. It also addresses the deficiencies of existing distributed monitoring methods, which rely on prior knowledge or data-driven criteria to divide sub-blocks, potentially severing cross-sub-block variable coupling relationships and failing to quantitatively evaluate the fault sensitivity of the features themselves. By introducing correlation analysis between score vectors and original process variables, using the sum of absolute correlation values ​​as a metric, this method selects score vectors highly sensitive to abnormal fluctuations and constructs a correlation score matrix. This method retains the comprehensive representational ability of the score vectors after PCA transformation for the original multivariate global coupling information, while eliminating the reliance on variance contribution rate ranking. It can accurately extract feature components carrying key fault information, effectively avoiding the omission of fault-sensitive information and eliminating irrelevant redundant interference.

[0021] 2. This invention uses the selected relevance score matrix as input features to train a Support Vector Data Description (SVDD) single-classification model. SVDD is not restricted by the prior assumption that the data must satisfy a Gaussian distribution or linear relationship. It uses a kernel function to nonlinearly map features to a high-dimensional space, and solves for the minimum hypersphere that encloses normal samples in the feature space to obtain a compact fault detection boundary. Compared with existing decentralized methods that directly model on a subset of original variables or traditional methods that use unfiltered full-dimensional score features, this invention significantly reduces feature redundancy and noise interference through pre-selection of relevance metrics. This makes the decision boundary of the SVDD model clearer, and the control limits more compact and accurate, thereby greatly improving the model's adaptability and robustness to complex nonlinear and non-Gaussian industrial process data.

[0022] 3. During online detection, this invention only requires calling the relevant projection matrix for mapping, resulting in low computational load and good real-time performance. Experimental results show that compared with traditional methods such as SM-SVDD, the CSM-SVDD method of this invention significantly improves the detection rate of various faults while maintaining the same number of features. The average detection rate increases from 59.65% of the traditional SM method and 69.73% of the traditional SM-SVDD method to 78.25% of the CSM-SVDD method of this invention. For extremely difficult-to-detect unknown faults (such as TE process fault 20), the detection rate is greatly improved from 35.25% to 88.12%, and the false alarm rate is low under normal operating conditions, proving that this method has outstanding technical advantages in early warning of weak faults and capture of complex unknown faults. Attached Figure Description

[0023] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the drawings without creative effort.

[0024] Figure 1 This is a flowchart of the fault detection method based on the correlation score matrix and support vector data described in this invention; Figure 2 This is a process flow diagram of the Tennessee-Eastman process in Embodiment 3 of the present invention; Figure 3 A comparison chart showing the changing trends of fault sensitivity between the score matrix (SM) extracted by the traditional PCA model and the score matrix (CSM) extracted by the method of this invention; Figure 4 The graph shows the detection results of the CSM-SVDD method of this invention and other methods for TE process fault 4. Figure 5 The graph shows the detection results of the CSM-SVDD method of this invention and other methods for TE process fault 20. Detailed Implementation

[0025] The embodiments of the technical solution of the present invention will now be described in detail with reference to the accompanying drawings. These embodiments are merely illustrative of the technical solution of the present invention and are therefore intended to limit the scope of protection of the present invention.

[0026] Unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains; the terminology used in this invention is for the purpose of describing particular embodiments only and is not intended to limit the invention; the terms “comprising” and “having”, and any variations thereof, in the specification, claims and foregoing description of the drawings are intended to cover non-exclusive inclusion.

[0027] In this invention, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least some embodiments of the invention. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this invention can be combined with other embodiments.

[0028] Example 1 This invention provides a fault detection method based on correlation score matrix and support vector data description. This method is applied to fault detection of complex industrial process data with nonlinear and non-Gaussian distribution characteristics, and includes the following steps: Step 1: Obtain historical data samples of the industrial process under normal operating conditions, construct the original dataset and standardize it to obtain a standard dataset with a mean of 0 and a variance of 1, and record the mean and variance parameters used in the standardization process, and use the standard dataset as the training set. Step 2: Calculate the covariance matrix corresponding to the standard dataset, and perform eigenvalue decomposition on the covariance matrix to obtain all eigenvalues ​​and the eigenvectors corresponding to each eigenvalue. The projection matrix is ​​constructed from all eigenvectors. The standard dataset is mapped through the projection matrix to obtain the full-dimensional score matrix. Step 3: Calculate the correlation between each score vector in the full-dimensional score matrix and each original process variable of the industrial process, and construct a correlation matrix between variables and score vectors; for each score vector, calculate the sum of the absolute values ​​of the correlation between the score vector and all process variables, and sort them in descending order according to the magnitude of the sum of absolute values; select the top-ranked score vectors as fault-sensitive score vectors according to a preset selection number, combine them to construct a correlation score matrix, and extract the feature vectors corresponding to the fault-sensitive score vectors from the projection matrix to construct a correlation projection matrix; Step 4: Use the relevant score matrix as the model input feature to construct a support vector data description model. Map the low-dimensional features to a high-dimensional space through nonlinear mapping. Solve for the minimum hypersphere that can wrap the normal training samples. Determine the center and radius of the minimum hypersphere. Set the radius of the hypersphere as the control limit for fault detection. Step 5: In the online detection phase, the current test data samples of the industrial system are collected in real time. The mean and variance parameters of the standardization process in Step 1 are used to standardize the test data samples. The relevant projection matrix in Step 3 is called to calculate the relevant score vector corresponding to the standardized test data samples. Step Six: Input the relevant score vector corresponding to the data sample to be tested into the trained support vector data description model, calculate the distance from the center of the smallest hypersphere after the relevant score vector is mapped to the high-dimensional feature space, and use this distance as the fault detection statistic; compare the fault detection statistic with the control limit. If the fault detection statistic is greater than the control limit, it is determined that the industrial system has failed; otherwise, it is determined to be operating normally.

[0029] Optionally, the standardization process described in step one is performed according to the following formula:

[0030] in, The original value of the variable. The mean of the variable. The standard deviation is the variable.

[0031] Optionally, the specific process of calculating the covariance matrix, decomposing the eigenvalues, and solving the full-dimensional score matrix in step two is as follows: Let the standard dataset be Calculate the covariance matrix , n The number of training set samples; for the covariance matrix Perform eigenvalue decomposition ,in It is a diagonal matrix composed of eigenvalues ​​sorted in descending order. The projection matrix is ​​composed of the corresponding feature vectors; the formula for calculating the full-dimensional score matrix is: , This is the full-dimensional score matrix obtained after mapping.

[0032] Optionally, the specific process of calculating the relevance, filtering the fault-sensitive score vectors, and constructing the relevance score matrix in step three is as follows: The Pearson correlation coefficient was used to calculate the first... Original process vectors With the score vectors Relevance A two-dimensional correlation matrix is ​​formed. ; Regarding the first For each score vector, calculate the sum of the absolute values ​​of its correlation with all process variables. ,in v The total number of industrial process variables; by The score vectors are sorted in descending order of value. A predetermined number of score vectors are selected, and the top-ranked vectors are used as fault-sensitive score vectors. These vectors are then combined to construct a relevant score matrix. Extract the feature vectors corresponding to the fault-sensitive score vectors from the projection matrix, and combine them to construct the relevant projection matrix. .

[0033] Optionally, the specific process of constructing and solving the support vector data description model in step four is as follows: With the relevant score matrix As training samples, a support vector data description model is constructed to optimize the objective, which aims to minimize the volume of the hypersphere. Normal samples are constrained to fall inside the hypersphere or on its surface. The optimization expression is as follows:

[0034] in, The radius of the hypersphere, The center of the hypersphere, The penalty coefficient is... As slack variables, The number of training samples. It is a nonlinear mapping function. For the th in the relevant score matrix One sample; By introducing Lagrange multipliers, the constrained optimization problem is transformed into a Lagrange dual optimization problem for solution. A Gaussian kernel function is chosen to replace the high-dimensional inner product operation; the expression for the Gaussian kernel function is: ,in Let the Gaussian kernel width parameter be denoted; solve the dual optimization problem to obtain the center of the hypersphere. With the radius of the hypersphere , radius As a fault detection and control limit.

[0035] Optionally, the formula for calculating the relevance score vector of the test data sample in step five is:

[0036] in, For the standardized test data, This represents the relevant score corresponding to the data to be tested.

[0037] Optionally, the logic for calculating fault detection statistics and determining faults in step six is ​​as follows: The formula for calculating the fault detection statistic of the sample to be tested is: ; The logic for determining whether a fault has occurred is as follows: if the fault detection statistic of the sample under test... If the condition is met, the current industrial system is determined to be faulty; otherwise, the system is determined to be in normal operating condition.

[0038] Example 2 like Figure 1 As shown, this invention provides a fault detection method based on relevance score matrix and support vector data description, comprising two stages: offline modeling and online detection. The specific steps are as follows: Step S1: Obtain a set of historical data samples of the industrial process under normal operating conditions, and construct the original process data matrix. ( v (Number of variables). Using the formula The dataset is standardized to obtain a standard dataset with a mean of 0 and a variance of 1.

[0039] Step S2: Calculate the covariance matrix of the standard dataset. ,right Perform eigenvalue decomposition to obtain the projection matrix consisting of all eigenvalues ​​and their corresponding eigenvectors. Mapping the standard dataset to the projection matrix The score matrix is ​​obtained. .

[0040] Step S3: Calculate each score vector Compared with the original process variables Relevance Obtain the relevance matrix .

[0041]

[0042] For the Given a score vector, calculate its score vector and all... v The sum of the absolute values ​​of the correlation of each process variable :

[0043] according to Sort the values ​​in descending order and select... The correlation score matrix is ​​formed by several sensitive score vectors with relatively large values. , its in The corresponding eigenvectors form the relevant projection matrix. .

[0044] Step S4: Using the relevant score matrix As training input, a support vector data description model is constructed. A kernel function is used to map the features to a high-dimensional space, and the optimization objective of minimizing the hypersphere volume is solved.

[0045] Introducing Lagrange multipliers and Construct the Lagrangian function, and the dual form of the above optimization problem is:

[0046] After solving the Lagrange duality problem, the minimum hypersphere center of the package normal sample is obtained. With the radius of the hypersphere (Support vectors located on the hypersphere) (distance to the center of the ball)

[0047] This radius As a control limit for fault detection.

[0048] Step S5: In the online testing phase, real-time acquisition of test data samples from the industrial system. The mean and variance from the offline modeling were used to standardize the sample.

[0049] Step S6: Call the projection matrix Calculate the relevant score vector of the test sample. .

[0050] Step S7: After calculating the correlation score vector of the test sample and mapping it to the high-dimensional feature space, transfer it to the center of the hypersphere. distance : (6) Will With the radius of the hypersphere Compare. If If the condition is met, the system is determined to be in a fault state; otherwise, the system is determined to be in a normal operating state.

[0051] Example 3 like Figure 2 As shown, the TE process simulates a continuously operating gas-liquid phase reaction system, including units such as a reactor, condenser, compressor, separator, and stripping tower. The process incorporates 41 measured variables and 11 operational variables. This invention collects 960 sets of data under normal operating conditions as a training set, and introduces test sets for various failure modes after the 160th sampling point.

[0052] 1. Offline modeling stage: 1.1 Obtaining data during normal TE process operation The standardization process is performed with a mean of 0 and a variance of 1.

[0053] 1.2 Calculate the covariance matrix and perform eigenvalue decomposition to obtain the full-dimensional score matrix. If the traditional PCA method (SM method) is used, the principal components are usually selected based on the 85% cumulative variance contribution rate. However, analysis reveals that the principal components, which contain most of the process data, may not necessarily show effective bias when a fault occurs.

[0054] 1.3. The relevant score matrix extraction strategy of this invention is adopted: the sum of the correlation coefficients between each score vector and 52 process variables is calculated. .according to The values ​​are sorted in descending order, and the same number of highly sensitive score vectors as in traditional methods are selected to form the score vectors. Observation of the selected score vector reveals that the score vector selected by the CSM method exhibits a significant trend deviation within the fault range, demonstrating extremely high sensitivity to faults. For example... Figure 3 As shown, Figure 3A comparison chart showing the changing trends of fault sensitivity between the score matrix (SM) extracted by the traditional PCA model and the score matrix (CSM) extracted by the method of this invention is presented.

[0055] 1.4. Input the CSM matrix into the SVDD model and set the penalty parameters ( ) and Gaussian kernel function parameters ( By solving the Lagrange dual problem, a minimum hypersphere encapsulating normal data is constructed, and the control limit radius is obtained. .

[0056] 2. Online testing phase: 2.1 In the operation of the TE process, it is assumed that fault IDV(4) (change in reactor cooling water inlet temperature, which is a step fault) or IDV(20) (an unknown fault that is extremely difficult to detect) is introduced.

[0057] 2.2 For the test sample, calculate its relevant score vector using the above logic and obtain the distance statistic. .

[0058] 2.3 The detection rate and false alarm rate for 21 types of faults in the TE process are shown in Table 1. With the same number of score vectors selected, the detection rate of the CSM method is generally better than that of the SM method, indicating that the score vectors selected based on relevance contain more fault information and can more effectively capture fault fluctuations, verifying its effectiveness in extracting sensitive features. The detection performance of SM-SVDD and CSM-SVDD is significantly better than linear methods. This is because SVDD can handle the nonlinear and non-Gaussian features in the TE process, mapping the data to a high-dimensional space through nonlinear mapping and then constructing a hypersphere boundary, thereby achieving more accurate identification of fault conditions. The CSM-SVDD method integrates the advantages of relevance feature extraction and SVDD modeling, achieving the highest detection rate in most fault modes. To more intuitively demonstrate the fault detection performance of different methods, Figure 4 and Figure 5 The fault detection results for fault 4 and fault 20 under the four methods are given respectively.

[0059] Table 1 Fault Detection Rate

[0060] 2.4. Regarding fault 4: Although the SM and CSM methods based on traditional linear feature selection respond in the early stages of the fault, the statistics frequently fall below the control limits afterward, resulting in serious false negatives, with detection rates of only 50.38% and 62.62%, respectively. In contrast, the CSM-SVDD method of this invention can quickly respond after the fault occurs and maintain its statistics above the control limits, significantly improving the detection rate to 85.25%.

[0061] 2.5 For fault 20: Traditional PCA and CSM methods almost failed, with a large number of fault samples being misclassified as normal, resulting in an extremely low detection rate (35.25%). In contrast, the CSM-SVDD method demonstrated significant advantages. Due to the accurate extraction of strongly correlated sensitive features and the construction of a nonlinear hypersphere boundary, the statistics clearly exceeded the control limits within the fault interval, achieving a detection rate as high as 88.12%, and exhibiting an extremely low false alarm rate within the normal operating condition interval.

[0062] In summary, this invention provides a fault detection method based on correlation score matrix and support vector data description, applicable to high-dimensional data status monitoring and early fault warning in complex industrial control systems such as chemical, power, and metallurgical industries. The method first performs principal component analysis on the acquired normal operation process data to obtain a full-dimensional score matrix. Next, it introduces the concept of variable coupling strength, calculating the correlation between each score vector and the original process variables. Based on the correlation, score vectors with high sensitivity to abnormal fluctuations are selected to construct a correlation score matrix. Then, the extracted correlation score matrix is ​​used as a feature input to the support vector data description model, constructing a minimum hypersphere boundary in the feature space that encloses normal samples, serving as the control limit for fault discrimination. Finally, the distance from the correlation score vector of the test sample to the center of the hypersphere is calculated online to achieve fault detection. This invention effectively overcomes the shortcomings of traditional principal component analysis, which relies solely on variance contribution rate to select principal components, easily retaining redundant information and ignoring fault-sensitive information. It also adapts to the nonlinearity and non-Gaussianity of industrial processes, significantly improving the fault detection rate and reducing the false alarm rate.

[0063] In some embodiments, the present invention provides a fault detection system based on correlation score matrix and support vector data description, for implementing the aforementioned fault detection method based on correlation score matrix and support vector data description, the system comprising: The data acquisition and standardization preprocessing module is used to collect normal historical data and online test data of industrial processes, and to complete the standardization process according to preset rules to generate training sets and standardized test datasets respectively. The covariance eigenvalue decomposition and score matrix solving module is used to calculate the covariance matrix of the standard dataset, perform eigenvalue decomposition on the covariance matrix to obtain the projection matrix, and combine the projection matrix mapping to solve the full-dimensional score matrix. The sensitive feature filtering module is used to calculate the correlation between the score vector and the original process variable, sort them in descending order according to the sum of the absolute values ​​of the correlation, filter the fault-sensitive score vectors according to the preset selection quantity, and construct the correlation score matrix and the correlation projection matrix. The SVDD offline modeling module is used to train a support vector data description model with the relevant score matrix as input, solve for the center and radius of the minimum hypersphere, and determine the fault detection control limits. The online feature calculation module is used to call the relevant projection matrix and calculate the relevant score vector corresponding to the standardized test data. The fault detection module is used to calculate the fault detection statistics of the sample under test, compare the statistics with the control limits, and output the operating status judgment result of the industrial system.

[0064] In some embodiments, the present invention provides a computer-readable storage medium storing computer instructions that are executed by a processor as described in any of the above embodiments regarding a fault detection method based on a correlation score matrix and support vector data.

[0065] Computer-readable storage media may take the form of any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples (not an exhaustive list) of readable storage media may include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0066] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0067] Embodiments of the present invention may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform the steps in the fault detection methods based on correlation score matrix and support vector data described in the "Exemplary Methods" section above this specification, according to various embodiments of the present invention.

[0068] The steps of the method of the present invention are not limited to the specific order described above, unless otherwise specifically stated. Furthermore, in some embodiments, the invention may also be implemented as a program recorded on a recording medium, the program comprising machine-readable instructions for implementing the method according to the invention. Therefore, the invention also covers recording media storing programs for performing the method according to the invention.

[0069] Although the invention has been described with reference to preferred embodiments, various modifications can be made and components can be replaced with equivalents without departing from the scope of the invention. In particular, the technical features mentioned in the various embodiments can be combined in any manner as long as there is no structural conflict. The invention is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

Claims

1. A fault detection method based on correlation score matrix and support vector data description, characterized in that, Includes the following steps: Step 1: Obtain historical data samples of the industrial process under normal operating conditions, construct the original dataset and standardize it to obtain a standard dataset with a mean of 0 and a variance of 1, and use the standard dataset as the training set; Step 2: Calculate the covariance matrix corresponding to the standard dataset, and perform eigenvalue decomposition on the covariance matrix to obtain all eigenvalues ​​and the eigenvectors corresponding to each eigenvalue. The projection matrix is ​​constructed from all eigenvectors. The standard dataset is mapped through the projection matrix to obtain the full-dimensional score matrix. Step 3: Calculate the correlation between each score vector in the full-dimensional score matrix and each original process variable of the industrial process, and construct a correlation matrix between variables and score vectors; for each score vector, calculate the sum of the absolute values ​​of the correlation between the score vector and all process variables, and sort them in descending order according to the magnitude of the sum of absolute values; select the top-ranked score vectors as fault-sensitive score vectors according to a preset selection number, combine them to construct a correlation score matrix, and extract the feature vectors corresponding to the fault-sensitive score vectors from the projection matrix to construct a correlation projection matrix; Step 4: Use the relevant score matrix as the model input feature to construct a support vector data description model. Map the low-dimensional features to a high-dimensional space through nonlinear mapping. Solve for the minimum hypersphere that can wrap the normal training samples. Determine the center and radius of the minimum hypersphere. Set the radius of the hypersphere as the control limit for fault detection. Step 5: In the online detection stage, the current test data samples of the industrial system are collected in real time, and the mean and variance parameters of the standardization process in Step 1 are used to standardize the test data samples. Call the relevant projection matrix from step three to calculate the relevant score vector corresponding to the standardized test data sample; Step Six: Input the relevant score vector corresponding to the data sample to be tested into the trained support vector data description model, calculate the distance from the center of the smallest hypersphere after the relevant score vector is mapped to the high-dimensional feature space, and use this distance as the fault detection statistic; compare the fault detection statistic with the control limit. If the fault detection statistic is greater than the control limit, it is determined that the industrial system has failed; otherwise, it is determined to be operating normally.

2. The fault detection method based on correlation score matrix and support vector data description according to claim 1, characterized in that, The standardization process described in step one is performed according to the following formula: in, The original value of the variable. The mean of the variable. The standard deviation is the variable.

3. The fault detection method based on correlation score matrix and support vector data description according to claim 2, characterized in that, The specific process of calculating the covariance matrix, decomposing the eigenvalues, and solving the full-dimensional score matrix in step two is as follows: Let the standard dataset be Calculate the covariance matrix , n The number of training set samples; for the covariance matrix Perform eigenvalue decomposition ,in It is a diagonal matrix composed of eigenvalues ​​sorted in descending order. The projection matrix is ​​composed of the corresponding feature vectors; the formula for calculating the full-dimensional score matrix is: , This is the full-dimensional score matrix obtained after mapping.

4. The fault detection method based on correlation score matrix and support vector data description according to claim 3, characterized in that, The specific process of calculating relevance, filtering fault-sensitive score vectors, and constructing the relevance score matrix in step three is as follows: The Pearson correlation coefficient was used to calculate the first... Original process vectors With the score vectors Relevance A two-dimensional correlation matrix is ​​formed. ; Regarding the first For each score vector, calculate the sum of the absolute values ​​of its correlation with all process variables. ,in v The total number of industrial process variables; by The score vectors are sorted in descending order of value. A predetermined number of score vectors are selected, and the top-ranked vectors are used as fault-sensitive score vectors. These vectors are then combined to construct a relevant score matrix. Extract the feature vectors corresponding to the fault-sensitive score vectors from the projection matrix, and combine them to construct the relevant projection matrix. .

5. The fault detection method based on correlation score matrix and support vector data description according to claim 4, characterized in that, The specific process of constructing and solving the support vector data description model described in step four is as follows: With the relevant score matrix As training samples, a support vector data description model is constructed to optimize the objective, which aims to minimize the volume of the hypersphere. Normal samples are constrained to fall inside the hypersphere or on its surface. The optimization expression is as follows: in, The radius of the hypersphere The center of the hypersphere, The penalty coefficient is... As slack variables, The number of training samples. It is a nonlinear mapping function. For the th in the relevant score matrix One sample; By introducing Lagrange multipliers, the constrained optimization problem is transformed into a Lagrange dual optimization problem for solution. A Gaussian kernel function is chosen to replace the high-dimensional inner product operation; the expression for the Gaussian kernel function is: ,in Let the Gaussian kernel width parameter be denoted; solve the dual optimization problem to obtain the center of the hypersphere. With the radius of the hypersphere , radius As a fault detection and control limit.

6. The fault detection method based on correlation score matrix and support vector data description according to claim 5, characterized in that, The formula for calculating the relevance score vector of the test data sample in step five is: in, For the standardized test data, This represents the relevant score corresponding to the data to be tested.

7. The fault detection method based on correlation score matrix and support vector data description according to claim 6, characterized in that, The logic for calculating fault detection statistics and determining faults in step six is ​​as follows: The formula for calculating the fault detection statistic of the sample to be tested is: ; The logic for determining whether a fault has occurred is as follows: if the fault detection statistic of the sample under test... If the condition is met, the current industrial system is determined to be faulty; otherwise, the system is determined to be in normal operating condition.

8. A fault detection system based on correlation score matrix and support vector data description, characterized in that, The system is used to implement the fault detection method based on correlation score matrix and support vector data as described in any one of claims 1-7, the system comprising: The data acquisition and standardization preprocessing module is used to collect normal historical data and online test data of industrial processes, and to complete the standardization process according to preset rules to generate training sets and standardized test datasets respectively. The covariance eigenvalue decomposition and score matrix solving module is used to calculate the covariance matrix of the standard dataset, perform eigenvalue decomposition on the covariance matrix to obtain the projection matrix, and combine the projection matrix mapping to solve the full-dimensional score matrix. The sensitive feature filtering module is used to calculate the correlation between the score vector and the original process variable, sort them in descending order according to the sum of the absolute values ​​of the correlation, filter the fault-sensitive score vectors according to the preset selection quantity, and construct the correlation score matrix and the correlation projection matrix. The SVDD offline modeling module is used to train a support vector data description model with the relevant score matrix as input, solve for the center and radius of the minimum hypersphere, and determine the fault detection control limits. The online feature calculation module is used to call the relevant projection matrix and calculate the relevant score vector corresponding to the standardized test data. The fault detection module is used to calculate the fault detection statistics of the sample under test, compare the statistics with the control limits, and output the operating status judgment result of the industrial system.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that are executed by a processor according to the fault detection method described in any one of claims 1-7 based on a correlation score matrix and support vector data.

10. A computer program product, characterized in that, The computer program product stores computer instructions, which are executed by a processor as described in any one of claims 1-7, for the fault detection method based on the correlation score matrix and support vector data.