Fault detection method, electronic device and readable storage medium based on improved KECA
Through the improved KECA method, the kernel matrix, Rayleigh entropy and CS statistics are used, combined with the FastICA algorithm, the problem of insufficient fault detection capabilities in complex industrial processes is solved, and more efficient fault detection and more stable fault detection rates are achieved.
Patent Information
- Application Number
- CN202110916951.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-10
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-08-10
AI Technical Summary
The existing technology has insufficient fault detection capabilities in complex industrial processes, resulting in stagnation of production processes, affecting product quality and safety, and the existing KECA method has low fault detection rate.
Using the improved KECA method, the feature extraction and fault detection capabilities are improved by determining the kernel matrix, Rayleigh entropy contribution size sorting, kernel space mapping and CS statistics, and combined with the FastICA algorithm.
It improves the accuracy and robustness of fault detection, enhances the stability of kernel function parameters, and improves the fault detection rate.
Smart Images

Figure CN113743476B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of fault detection, and in particular to a fault detection method based on improved KECA, an electronic device and a readable storage medium. Background Art
[0002] Complex industrial production processes are complex and often accompanied by extreme conditions such as high temperature and high pressure. Raw materials and products are flammable, explosive, toxic and harmful, and production equipment is large-scale and continuous. If a minor fault occurs in the equipment of an industrial process during the production process and is not discovered and eliminated in time, the production process may stagnate, affecting product quality, failing to protect people's life safety and social property safety, and exacerbating environmental pollution and hazards. In recent years, multivariate statistical process monitoring has been formed in the monitoring of complex industrial processes. Its basic idea is to summarize the information carried by high-dimensional data by constructing a set of unrelated latent variables with low dimensions. Common multivariate statistical process monitoring methods include principal component analysis (PCA), partial least squares (PLS), independent principal component analysis (ICA), and KPCA for processing nonlinear data. Among them, the KPCA algorithm is a nonlinear enhancement algorithm of PCA. By introducing the idea of kernel theory, it realizes the problem of using PCA to process nonlinear data. KECA is obtained by introducing information to transform KPCA. It mainly uses the information entropy of the original input data as the main indicator, rather than the variance of the feature data. In the process of feature extraction, the KECA algorithm focuses on the maximum contribution rate of the eigenvalue and the corresponding eigenvector to the Rayleigh entropy of the original feature data set.
[0003] The above contents are only used to assist in understanding the technical solution of the present invention and do not constitute an admission that the above contents are prior art. Summary of the invention
[0004] The main purpose of the present invention is to provide a fault detection method based on improved KECA, aiming to improve the fault detection capability.
[0005] In order to achieve the above object, the present invention provides a fault detection method based on improved KECA, and the fault detection method based on improved KECA comprises the following steps:
[0006] Determine the kernel matrix based on the data set and the indefinite kernel function;
[0007] Sort the elements in the kernel matrix according to the Rayleigh entropy contribution and determine the score matrix;
[0008] Determine a kernel space mapping of the score matrix, and obtain a kernel space vector according to the kernel space matrix;
[0009] Determine a CS statistic according to the kernel space vector and the kernel space average vector;
[0010] When the CS statistic exceeds the control limit, a fault occurrence prompt message is sent.
[0011] Furthermore, after the step of determining the kernel space mapping of the score matrix, the step further includes:
[0012] performing whitening processing on the kernel space mapping;
[0013] A rotation matrix is determined, and a kernel space vector is determined according to the rotation matrix and the kernel space mapping after whitening processing.
[0014] Furthermore, the step of determining the kernel matrix according to the data set and the indefinite kernel function includes:
[0015] The data were standardized according to the mean and variance in the centrifugal model of the KECA algorithm;
[0016] The kernel matrix is determined based on the standardized data and the indefinite kernel function.
[0017] Furthermore, before the step of sorting the elements in the kernel matrix according to the Rayleigh entropy contribution and determining the score matrix, the step further includes:
[0018] Determining the eigenvalues and eigenvectors of the kernel matrix;
[0019] Determine a density function corresponding to the data set, and determine a Rayleigh entropy calculation formula for the data set based on the density function;
[0020] Determining an approximate value of the density function according to an indefinite kernel function, wherein the approximate value of the density function can be obtained according to the indefinite kernel function;
[0021] Determining the Rayleigh entropy estimation value calculation formula according to the approximate value of the density function;
[0022] The Rayleigh entropy estimation value is determined according to the eigenvalue, the eigenvector and the Rayleigh entropy estimation value calculation formula.
[0023] Furthermore, the step of determining the CS statistic according to the kernel space vector and the kernel space average vector includes:
[0024] Determine the cosine value of the kernel space vector and the kernel space mean vector;
[0025] A corresponding CS statistic is determined according to the cosine value.
[0026] Furthermore, before the step of sending a fault occurrence prompt message when the CS statistic exceeds the control limit, the step further includes:
[0027] Standardize the test data;
[0028] Establish a KECA model based on the indefinite kernel function, kernel parameters and modeling data;
[0029] Determine the kernel matrix according to the standardized test data and the indefinite kernel function;
[0030] Determine the Rayleigh entropy value according to the eigenvalue and eigenvector corresponding to the kernel matrix;
[0031] According to a preset number of elements contributing to the Rayleigh entropy value, and determining the kernel space mapping data of the elements;
[0032] Performing whitening processing on the kernel space mapping data, and performing fast independent principal component analysis on the whitening processing result to obtain a kernel space vector;
[0033] Determine a statistic by determining each of the kernel space vectors and the kernel space average vector;
[0034] The control limit corresponding to the statistic is determined by a preset significance level.
[0035] Furthermore, the step of performing fast independent principal component analysis on the whitening processing result to obtain the kernel space vector includes:
[0036] The kernel space vector is obtained by extracting features from the whitening results through fast independent principal component analysis.
[0037] In order to achieve the above-mentioned objectives, the present invention also provides an electronic device, comprising a memory, a processor, and a fault detection program based on an improved KECA stored in the memory and executable on the processor, wherein the fault detection program based on the improved KECA, when executed by the processor, implements the steps of the fault detection method based on the improved KECA described above.
[0038] In order to achieve the above-mentioned objectives, the present invention also provides a readable storage medium, on which a fault detection program based on an improved KECA is stored. When the fault detection program based on the improved KECA is executed by a processor, the steps of any of the above-mentioned fault detection methods based on the improved KECA are implemented.
[0039] In the technical solution of the present invention, a kernel matrix is determined according to a data set and an indefinite kernel function; elements in the kernel matrix are sorted according to the Rayleigh entropy contribution, and a score matrix is determined; the kernel space mapping of the score matrix is determined, and a kernel space vector is obtained according to the kernel space matrix; a CS statistic is determined according to the kernel space vector and the kernel space average vector; when the CS statistic exceeds the control limit, a fault occurrence prompt message is sent. In this way, by selecting eigenvalues and eigenvectors according to the Rayleigh entropy value in the fault detection process to reduce the amount of calculation, selecting kernel function decomposition based on the indefinite kernel function method to improve the robustness of the fault detection rate to the kernel function parameters, introducing the FastICA algorithm in the feature extraction stage to maximize the feature extraction capability, and using the CS statistic to better express the similarity between different probability density distributions, the fault detection capability is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is a schematic diagram of the device structure of the hardware operating environment involved in the embodiment of the present invention;
[0041] Figure 2 It is a flow chart of an embodiment of a fault detection method based on improved KECA of the present invention;
[0042] Figure 3 This is a detailed flow chart of step S10 in the fault detection method based on improved KECA of the present invention.
[0043] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0044] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0045] The main technical scheme of the present invention is:
[0046] Determine the kernel matrix based on the data set and the indefinite kernel function;
[0047] Sort the elements in the kernel matrix according to the Rayleigh entropy contribution and determine the score matrix;
[0048] Determine a kernel space mapping of the score matrix, and obtain a kernel space vector according to the kernel space matrix;
[0049] Determine a CS statistic according to the kernel space vector and the kernel space average vector;
[0050] When the CS statistic exceeds the control limit, a fault occurrence prompt message is sent.
[0051] In the related art, KPCA is transformed into KECA by introducing information entropy. KECA mainly uses the information entropy of the original input data as the main indicator, rather than the variance of the feature data. The KECA algorithm focuses on the maximum contribution rate of the eigenvalue and the corresponding eigenvector to the Rayleigh entropy of the original feature data set during the feature extraction process.
[0052] In the technical solution of the present invention, a kernel matrix is determined according to a data set and an indefinite kernel function; elements in the kernel matrix are sorted according to the Rayleigh entropy contribution, and a score matrix is determined; the kernel space mapping of the score matrix is determined, and a kernel space vector is obtained according to the kernel space matrix; a CS statistic is determined according to the kernel space vector and the kernel space average vector; when the CS statistic exceeds the control limit, a fault occurrence prompt message is sent. In this way, by selecting eigenvalues and eigenvectors according to the Rayleigh entropy value in the fault detection process to reduce the amount of calculation, selecting kernel function decomposition based on the indefinite kernel function method to improve the robustness of the fault detection rate to the kernel function parameters, introducing the FastICA algorithm in the feature extraction stage to maximize the feature extraction capability, and using the CS statistic to better express the similarity between different probability density distributions, the fault detection capability is improved.
[0053] like Figure 1 As shown, Figure 1 It is a schematic diagram of the hardware operating environment of the terminal involved in the embodiment of the present invention.
[0054] like Figure 1 As shown, the terminal may include: a processor 1001, such as a CPU, a network interface 1004, a user interface 1003, a memory 1005, and a communication bus 1002. The communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a touch screen and / or buttons, etc., and the user interface 1003 may optionally include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a memory (non-volatile memory), such as a disk storage. The memory 1005 may optionally be a storage device independent of the aforementioned processor 1001.
[0055] Those skilled in the art will understand that Figure 1 The structure of the terminal shown in the figure does not constitute a limitation on the terminal, and may include more or less components than shown in the figure, or combine certain components, or arrange the components differently.
[0056] like Figure 1 As shown, the memory 1005 as a computer storage medium may include an operating system, a network communication module, a user interface module, and a fault detection program based on the improved KECA.
[0057] exist Figure 1 In the terminal shown, the network interface 1004 is mainly used to connect to the backend server and perform data communication with the backend server; the user interface 1003 is mainly used to connect to the client (user end) and perform data communication with the client; and the processor 1001 can be used to call the improved KECA-based fault detection program stored in the memory 1005 and perform the following operations:
[0058] Determine the kernel matrix based on the data set and the indefinite kernel function;
[0059] Sort the elements in the kernel matrix according to the Rayleigh entropy contribution and determine the score matrix;
[0060] Determine a kernel space mapping of the score matrix, and obtain a kernel space vector according to the kernel space matrix;
[0061] Determine a CS statistic according to the kernel space vector and the kernel space average vector;
[0062] When the CS statistic exceeds the control limit, a fault occurrence prompt message is sent.
[0063] Further, the processor 1001 may call the fault detection program based on the improved KECA stored in the memory 1005, and further perform the following operations:
[0064] performing whitening processing on the kernel space mapping;
[0065] A rotation matrix is determined, and a kernel space vector is determined according to the rotation matrix and the kernel space mapping after whitening processing.
[0066] Further, the processor 1001 may call the fault detection program based on the improved KECA stored in the memory 1005, and further perform the following operations:
[0067] The data were standardized according to the mean and variance in the centrifugal model of the KECA algorithm;
[0068] The kernel matrix is determined based on the standardized data and the indefinite kernel function.
[0069] Further, the processor 1001 may call the fault detection program based on the improved KECA stored in the memory 1005, and further perform the following operations:
[0070] Determining the eigenvalues and eigenvectors of the kernel matrix;
[0071] Determine a density function corresponding to the data set, and determine a Rayleigh entropy calculation formula for the data set based on the density function;
[0072] Determining an approximate value of the density function according to an indefinite kernel function, wherein the approximate value of the density function can be obtained according to the indefinite kernel function;
[0073] Determining the Rayleigh entropy estimation value calculation formula according to the approximate value of the density function;
[0074] The Rayleigh entropy estimation value is determined according to the eigenvalue, the eigenvector and the Rayleigh entropy estimation value calculation formula.
[0075] Further, the processor 1001 may call the fault detection program based on the improved KECA stored in the memory 1005, and further perform the following operations:
[0076] Determine the cosine value of the kernel space vector and the kernel space mean vector;
[0077] A corresponding CS statistic is determined according to the cosine value.
[0078] Further, the processor 1001 may call the fault detection program based on the improved KECA stored in the memory 1005, and further perform the following operations:
[0079] Standardize the test data;
[0080] Establish a KECA model based on the indefinite kernel function, kernel parameters and modeling data;
[0081] Determine the kernel matrix according to the standardized test data and the indefinite kernel function;
[0082] Determine the Rayleigh entropy value according to the eigenvalue and eigenvector corresponding to the kernel matrix;
[0083] According to a preset number of elements contributing to the Rayleigh entropy value, and determining the kernel space mapping data of the elements;
[0084] Performing whitening processing on the kernel space mapping data, and performing fast independent principal component analysis on the whitening processing result to obtain a kernel space vector;
[0085] Determine a statistic by determining each of the kernel space vectors and the kernel space average vector;
[0086] The control limit corresponding to the statistic is determined by a preset significance level.
[0087] Further, the processor 1001 may call the fault detection program based on the improved KECA stored in the memory 1005, and further perform the following operations:
[0088] The kernel space vector is obtained by extracting features from the whitening results through fast independent principal component analysis.
[0089] The KECA method in the related art selects the principal component to realize data feature extraction by calculating the contribution of the Rayleigh entropy value of the input space. Its basic mathematical principle is as follows: define the data set D = x1, x2, ... x n , the probability density function is p(x), and the Rayleigh entropy calculation formula of the data set can be obtained as follows:
[0090] H(p)=-lg∫p 2 (x)dx (1.1)
[0091] Since the logarithmic function is a monotonic function, we can calculate V(p) = ∫p 2 (x)dx can also be expressed as V in statistics. p =ε p (p), where ε p (p) represents the expectation of the density function p(x). The approximate value of p(x) can be obtained by the Parzen window density estimator The calculation formula is as follows:
[0092]
[0093] Among them, K σ (x,x t ) becomes a Parzen window, which is x t The kernel density function is centered and the width is determined by σ. In the process of applying the Parzen window, the radial basis kernel function (RBF) is often selected for data mapping. The calculation formula of the radial basis kernel function is as follows:
[0094]
[0095] The kernel density function K σ (x,x t ) and the density function approximation By approximating the sample mean of the expectation operator, we can get an estimate of V(p)
[0096]
[0097] K is an N×N kernel matrix, and all elements in K are calculated by (t, t') σ (x,x t ) is composed of; 1 is an N×1 vector whose elements are all 1. From the above formula, it can be seen that the Rayleigh entropy estimate obtained based on the sample is completely contained in the corresponding elements of the kernel matrix, so the eigenvalues and eigenvectors of the kernel matrix can be used to represent the entropy estimate of the Rayleigh entropy, where the kernel matrix can be characteristically decomposed into K=Φ T Φ=EDE T, where D is the eigenvalue matrix D = diag(λ1,λ2,...,λ N ), E is the eigenvector matrix E=(e1,e2,...,e N ), so The calculation formula can be shown as follows:
[0098]
[0099] When reducing the data dimension, the first k principal components with the largest entropy contribution are selected according to the entropy contribution rate. The data matrix composed of the selected principal components is the projected data Φ eca , as shown below:
[0100]
[0101] Among them, with Φ eca The associated entropy estimate The calculation formula is as follows:
[0102]
[0103] in, So Φ eca The calculation formula is as follows:
[0104]
[0105] The solution to the minimum value of equation (1.8) is given by the following equation, where Ψ j Corresponding to the jth largest term in equation (1.5):
[0106]
[0107] The test data outside the training sample is mapped to U by the KECA algorithm. k Φ' eca , Φ' eca The calculation formula is as follows:
[0108]
[0109] From the above, we can see that the KECA algorithm uses the Parzen window probability density to calculate the Rayleigh entropy estimate. Therefore, the choice of kernel function has an important impact on the final result. Common kernel functions include semi-definite kernel functions and indefinite kernel functions, among which the most commonly used is the semi-definite kernel function. However, the kernel function is not necessarily a positive definite kernel function. The Sigmoid kernel in the neural network is a common kernel function, which is non-positive definite under certain conditions. In the experimental process, it is found that except for some known kernel functions, it is difficult to detect its positive definiteness. In addition, recent studies have found that the positive definiteness of the kernel function itself has considerable limitations, and the theoretically optimal Parzen window also has an indefinite nature. Therefore, it is proposed to decompose the kernel function based on the indefinite kernel function method to improve the robustness of the fault detection rate to the kernel function parameters.
[0110] For faults, the effective fault detection rate of KECA's SPE method is very low, while the principal component selected by the KECA method in the present invention has an angular structure characteristic and adopts the CS (Cauchy-Schwarz) statistic, which corresponds to the angular cosine value between vectors in the kernel feature space, and can better express the similarity between different probability density distributions.
[0111] like Figure 2 As shown, in one embodiment of the present invention, the fault detection method based on improved KECA includes the following steps:
[0112] Step S10, determining a kernel matrix according to the data set and the indefinite kernel function;
[0113] In this embodiment, since different kernel functions usually perform inner product calculations in different feature spaces, when the kernel function is a semi-positive definite function, the inner product is calculated in a reproducing kernel Hilbert space (RKHS), and the Hilbert space has the following properties:
[0114] (1) Hilbert space is a linear space in the field of rational numbers;
[0115] (2) When the inner product number domain is Hilbert space, it has bilinearization, that is, it has the following properties:
[0116] <x,y> =<y,x> (1.11)
[0117] <ax+by,z> =a<x,z> +<x,z> +b<y,z>
[0118] Among them, x, y, z∈H, a, b∈R.
[0119] (3) When x≠0, we have<x,x> >0.
[0120] When the kernel function is an indefinite kernel function, the inner product will be obtained in the reproducing kernel Krein space (RKKS). The first two properties of the kernel Krein space are the same as those of the kernel Hilbert reconstruction space. In addition, the spatial properties of the kernel Krein space also include allowing direct sum decomposition.
[0121]
[0122] Among them, K + <.,.>, K-<.,.> belongs to the Krein space and satisfies<x,y> = 0, x∈K + <.,.>, y∈K-<.,.>.
[0123] From the above analysis, it can be seen that the inner product and norm in the Klein space no longer have to have positive definite properties. The function related to this theory is called an indefinite kernel function, so the Epanechnikov kernel function is used as a high-dimensional feature mapping function, as shown in the following formula:
[0124]
[0125] Therefore, this embodiment determines the corresponding kernel matrix according to the data set and the indefinite kernel function to improve the robustness of the fault detection rate to the kernel function parameters.
[0126] Step S20, sorting the elements in the kernel matrix according to the Rayleigh entropy contribution, and determining the score matrix;
[0127] In this embodiment, the KECA method is adopted to select the principal component according to the size of the Rayleigh entropy value, thereby reducing the loss of information in the process of data dimensionality reduction. This not only achieves a small number of principal components, but also ensures that the data after dimensionality reduction still retains more than 99% of the information entropy value of the original data in the kernel feature space.
[0128] Step S30, determining the kernel space mapping of the score matrix, and obtaining the kernel space vector according to the kernel space matrix;
[0129] In this embodiment, the score matrix is determined according to a kernel space mapping in a pre-generated algorithm centrifugal model, and the kernel space mapping is whitened to determine a corresponding kernel space vector.
[0130] Step S40, determining a CS statistic according to the kernel space vector and the kernel space average vector;
[0131] In this embodiment, the SPE statistic is an electrical distance-based detection indicator, and the SPE statistic is defined as:
[0132]
[0133] Where t is the sampling time, Among them, the row vector with the larger binorm in W is d. d , the rest is W e , then B d =P -1 W d , The SPE statistic describes the extent to which the correlation between variables has been changed, indicating an abnormal process.
[0134] The kernel entropy component analysis algorithm usually obtains data sets with different angular structures, where the clustering results of different data sets are more or less distributed in different angular directions relative to the origin of the kernel feature space. The CS divergence measure between data density functions corresponds to the cosine measure of the angle between the average vectors in the kernel feature space. For the normal training data set X∈D∈R d×N , where d is the measurement parameter, N is the number of samples, and the total kernel space mean vector m of the data is as follows:
[0135]
[0136] Where i = 1, ..., N, Φ eca (x i ) is the entropy component score vector in the kernel space, that is, x i In Φ eca (x i ) can be expressed as m i , which is m i =Φ eca (x i ). For the current i-th sampling point x i , define the CS statistic kernel space score vector m of the current point in the unkernel feature space i The cosine value of the angle between the vector m and the total kernel space average vector m is calculated as follows:
[0137]
[0138] The more similar the data are, the closer the CS statistic is to 1; the greater the difference between the data, the closer the CS statistic is to 0. In order to conform to the habit of process monitoring, the CS statistic is defined as follows:
[0139]
[0140] The CS statistic of the data outside the training sample is also the cosine of the angle with the total kernel space mean vector m:
[0141]
[0142] Among them, m i'is the data outside the training sample x' i In Φ' eca The corresponding vector in x' i Usually it is sampling data for online monitoring.
[0143] Step S50: When the CS statistic exceeds the control limit, a fault occurrence prompt message is sent.
[0144] In this embodiment, it is determined whether the statistical control limit under normal working conditions is exceeded by calculating the current respective statistical quantities. If the control limit is exceeded, information indicating that a fault has been detected is output to prompt the user.
[0145] In summary, this embodiment determines the kernel matrix according to the data set and the indefinite kernel function; sorts the elements in the kernel matrix according to the Rayleigh entropy contribution, and determines the score matrix; determines the kernel space mapping of the score matrix, and obtains the kernel space vector according to the kernel space matrix; determines the CS statistic according to the kernel space vector and the kernel space average vector; when the CS statistic exceeds the control limit, sends a fault occurrence prompt message. In this way, by selecting eigenvalues and eigenvectors according to the Rayleigh entropy value in the fault detection process to reduce the amount of calculation, selecting the kernel function decomposition based on the indefinite kernel function method to improve the robustness of the fault detection rate to the kernel function parameters, introducing the FastICA algorithm in the feature extraction stage to maximize the feature extraction capability, and using the CS statistic to better express the similarity between different probability density distributions, the fault detection capability is improved.
[0146] Furthermore, in one embodiment of the present invention, after the step of determining the kernel space mapping of the score matrix, the step further includes:
[0147] performing whitening processing on the kernel space mapping;
[0148] A rotation matrix is determined, and a kernel space vector is determined according to the rotation matrix and the kernel space mapping after whitening processing.
[0149] In this embodiment, the KECA algorithm is combined with the ICA theory. The purpose of ICA is to find the unknown mixing matrix A from the mixed signal X and realize the blind source separation of the source signal. Assume that there are m unknown source signals s i (i=1,2,...,m) to form a column vector matrix s=[s1,s2,...,s m ] T The mixing matrix is an l×m matrix X=[x1,x2,...,x l ] is the observed signal vector and X = A·S. The purpose of ICA is to find a separation matrix W that satisfies and It is the estimated value of the independent source signal S separated from the mixed signal. 1 / 2 E T ), the FastICA algorithm was introduced to improve the independence between components to obtain the maximum information potential, that is, the entropy value in formula (1.7). According to the basic principle of ICA, the data is whitened as shown in the following formula:
[0150]
[0151] It is necessary to find a rotation matrix W so that the elements in the independent element eigenvector S estimated by the following formula are independent of each other:
[0152] S=WX=B T Z (2.2)
[0153] Where W = B T P T ,B=[b1,b2,...,b N ] is an orthogonal decomposition matrix, that is, it satisfies the condition BB T =I N .
[0154] Using the method based on fast ICA theory, we get w k+1 The iteration formula is:
[0155]
[0156]
[0157] w k+1 =w″ k+1 / ||w″ k+1 || (2.5)
[0158] Among them, E(·) is the mathematical expectation, function g is the derivative of G,
[0159] The core goal of the present invention is to seek to represent as much information as possible with as few components as possible, so w is normalized and orthogonalized after iteration:
[0160]
[0161] w k+1 =w k+1 / ||w k+1 || (2.7)
[0162] like Figure 3 As shown, further, in one embodiment of the present invention, the step of determining the kernel matrix according to the data set and the indefinite kernel function includes:
[0163] Step S11, standardizing the data according to the mean and variance in the centrifugal model of the KECA algorithm;
[0164] Step S12, determining a kernel matrix according to the standardized data and the indefinite kernel function.
[0165] In this embodiment, the modeling data is used to perform KECA modeling in advance to obtain the KECA algorithm centrifugal model. After receiving the data set, the data in the data set is standardized by the mean and variance in the KECA algorithm centrifugal model. The kernel matrix is determined based on the standardized data and the indefinite kernel function selected in advance to determine the corresponding kernel space vector.
[0166] Furthermore, in one embodiment of the present invention, before the step of sorting the elements in the kernel matrix according to the Rayleigh entropy contribution and determining the score matrix, the step further includes:
[0167] Determining the eigenvalues and eigenvectors of the kernel matrix;
[0168] Determine a density function corresponding to the data set, and determine a Rayleigh entropy calculation formula for the data set based on the density function;
[0169] Determining an approximate value of the density function according to an indefinite kernel function, wherein the approximate value of the density function can be obtained according to the indefinite kernel function;
[0170] Determining the Rayleigh entropy estimation value calculation formula according to the approximate value of the density function;
[0171] The Rayleigh entropy estimation value is determined according to the eigenvalue, the eigenvector and the Rayleigh entropy estimation value calculation formula.
[0172] In this embodiment, it can be known from formula (1.4) that the Rayleigh entropy estimate corresponding to the data set is completely contained in the corresponding elements of the kernel matrix, so the eigenvalues and eigenvectors of the kernel matrix can be used to represent the entropy estimate of the Rayleigh entropy, where the kernel matrix can be characteristically decomposed into K = Φ T Φ=EDE T , where D is the eigenvalue matrix D = diag(λ1,λ2,...,λ N ), E is the eigenvector matrix E=(e1,e2,...,e N ). Determine the density function p(x) corresponding to the data set, and determine the Rayleigh entropy calculation formula of the data set based on the density function p(x), that is, H(p) = -lg∫p 2 (x)dx. Determine the approximate value of the density function p(x) according to the indefinite kernel function The calculation formula of the approximate value of the density function p(x) is: Among them, K σ (x,x t ) is an indefinite kernel function, so the density function can be determined based on the indefinite kernel function. The Rayleigh entropy estimate The calculation formula is:
[0173]
[0174] This includes an approximation of the density function Therefore, we can use the approximate value of the density function The Rayleigh entropy estimation value calculation formula is determined. Since the eigenvalues and eigenvectors of the kernel matrix can be used to represent the entropy estimation value of the Rayleigh entropy, the kernel matrix can be characteristically decomposed into K = Φ T Φ=EDE T , where D is the eigenvalue matrix D = diag(λ1,λ2,...,λ N ), E is the eigenvector matrix E=(e1,e2,...,e N ), so The calculation formula can be shown as follows:
[0175]
[0176] Therefore, the Rayleigh entropy estimation value may be determined according to the eigenvalue, the eigenvector and the Rayleigh entropy estimation value calculation formula.
[0177] Furthermore, in one embodiment of the present invention, the step of determining the CS statistic according to the kernel space vector and the kernel space average vector includes:
[0178] Determine the cosine value of the kernel space vector and the kernel space mean vector;
[0179] A corresponding CS statistic is determined according to the cosine value.
[0180] In this embodiment, the CS divergence measure between data density functions corresponds to the cosine measure of the angle between the average vectors in the kernel feature space. d×N , where d is the measurement parameter, N is the number of samples, and the total kernel space mean vector m of the data is as follows:
[0181]
[0182] Where i = 1, ..., N, Φ eca (x i ) is the entropy component score vector in the kernel space, that is, x i In Φ eca (x i ) can be expressed as mi , which is m i =Φ eca (x i ). For the current i-th sampling point x i , define the CS statistic kernel space score vector m of the current point in the unkernel feature space i The cosine value of the angle between the vector m and the total kernel space average vector m is calculated as follows:
[0183]
[0184] Since the principal components selected by the KECA method have angular structural characteristics, the CS statistic can better represent the similarity between different probability density distributions.
[0185] Furthermore, in an embodiment of the present invention, before the step of sending a fault occurrence prompt message when the CS statistic exceeds the control limit, the step further includes:
[0186] Standardize the test data;
[0187] Establish a KECA model based on the indefinite kernel function, kernel parameters and modeling data;
[0188] Determine the kernel matrix according to the standardized test data and the indefinite kernel function;
[0189] Determine the Rayleigh entropy value according to the eigenvalue and eigenvector corresponding to the kernel matrix;
[0190] According to a preset number of elements contributing to the Rayleigh entropy value, and determining the kernel space mapping data of the elements;
[0191] Performing whitening processing on the kernel space mapping data, and performing fast independent principal component analysis on the whitening processing result to obtain a kernel space vector;
[0192] Determine a statistic by determining each of the kernel space vectors and the kernel space average vector;
[0193] The control limit corresponding to the statistic is determined by a preset significance level.
[0194] In this embodiment, the test data is first standardized. According to a preset kernel function, such as the Epanechnikov kernel function, and kernel parameters, the modeling data is used to perform modeling to obtain a KECA algorithm centrifugal model, so that the data set can be processed by the KECA algorithm centrifugal model when the data set is acquired online. The kernel matrix is determined according to the standardized test function and the preset indefinite kernel function. The eigenvalues and eigenvectors of the kernel matrix are determined, and the corresponding Rayleigh entropy values are determined according to the eigenvalues and the eigenvectors. The elements in the kernel matrix are sorted according to their contribution to the Rayleigh entropy value, and a preset number of the elements are selected therefrom. The kernel space mapping data Φ of the preset number of the elements is determined. eca , for the low-dimensional embedding D 1 / 2 E T That is Φ eca After whitening, the whitening result Φ ecanew Perform fast independent component analysis (FastICA) to obtain the kernel space vector m i . Determine each kernel space score vector m i The cosine of the angle with the total kernel space mean vector m, i.e., the CS statistic and the SPE statistic. By setting the preset significance level δ, find the points occupying the 1-δ region of the estimated density function to determine the control limits of each statistic.
[0195] Furthermore, in one embodiment of the present invention, the step of performing fast independent principal component analysis on the whitening processing result to obtain the kernel space vector includes:
[0196] The kernel space vector is obtained by extracting features from the whitening results through fast independent principal component analysis.
[0197] In this embodiment, the KECA algorithm is combined with the ICA theory. The purpose of ICA is to find the unknown mixing matrix A from the mixed signal X and realize the blind source separation of the source signal. Assume that there are m unknown source signals s i (i=1,2,...,m) to form a column vector matrix s=[s1,s2,...,s m ] T The mixing matrix is an l×m matrix X=[x1,x2,...,x l ] is the observed signal vector and X = A·S. The purpose of ICA is to find a separation matrix W that satisfies and It is the estimated value of the independent source signal S separated from the mixed signal. 1 / 2 E T) was then introduced, namely the FastICA algorithm, which is the fast independent principal component analysis method, to improve the independence between components in order to obtain the maximum information potential, namely the entropy value in formula (1.7).
[0198] In order to achieve the above-mentioned objectives, the present invention also provides an electronic device, comprising a memory, a processor, and a fault detection program based on an improved KECA stored in the memory and executable on the processor, wherein the fault detection program based on the improved KECA, when executed by the processor, implements the steps of the fault detection method based on the improved KECA described above.
[0199] In order to achieve the above-mentioned objectives, the present invention also provides a readable storage medium, on which a fault detection program based on an improved KECA is stored. When the fault detection program based on the improved KECA is executed by a processor, the steps of any of the above-mentioned fault detection methods based on the improved KECA are implemented.
[0200] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0201] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a number of instructions for a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present invention.
[0202] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A fault detection method based on improved KECA, characterized in that: The fault detection method based on improved KECA comprises the following steps: Determine the kernel matrix based on the data set and the indefinite kernel function; Sort the elements in the kernel matrix according to the Rayleigh entropy contribution and determine the score matrix; Determine a kernel space mapping of the score matrix, and perform whitening processing on the kernel space mapping; Determine a rotation matrix, and determine a kernel space vector according to the rotation matrix and the kernel space mapping after whitening processing; Determine a CS statistic according to the kernel space vector and the kernel space average vector; Determine the control limit corresponding to the CS statistic: standardize the test data, establish a KECA model according to the indefinite kernel function, kernel parameters and modeling data, redetermine the kernel matrix according to the standardized test data and the indefinite kernel function, determine the Rayleigh entropy value according to the eigenvalues and eigenvectors corresponding to the redetermined kernel matrix, sort each element in the redetermined kernel matrix according to its contribution to the Rayleigh entropy value, select a preset number of elements therefrom, determine the kernel space mapping data of the preset number of elements, whiten the kernel space mapping data, and perform a fast independent principal component analysis on the whitening result to obtain a new kernel space vector, determine the statistic according to each of the new kernel space vectors and the new kernel space average vector, and find the points occupying the 1-δ region of the density function through a preset significance level δ to determine the control limit corresponding to the CS statistic; When the CS statistic exceeds the control limit, a fault occurrence prompt message is sent, wherein the indefinite kernel function performs inner product calculation in a kernel Krein reconstruction space, the kernel Krein reconstruction space is a linear space in the field of rational numbers, when the inner product belongs to the kernel Krein reconstruction space, it has a bilinear property, and the kernel Krein reconstruction space allows direct sum decomposition.
2. The fault detection method based on improved KECA according to claim 1, characterized in that: The step of determining the kernel matrix according to the data set and the indefinite kernel function comprises: The data were standardized according to the mean and variance in the centrifugal model of the KECA algorithm; The kernel matrix is determined based on the standardized data and the indefinite kernel function.
3. The fault detection method based on improved KECA according to claim 1, characterized in that: The step of determining the CS statistic according to the kernel space vector and the kernel space average vector comprises: Determine the cosine value of the kernel space vector and the kernel space mean vector; A corresponding CS statistic is determined according to the cosine value.
4. The fault detection method based on improved KECA according to claim 1, characterized in that: The step of performing fast independent principal component analysis on the whitening processing result to obtain a new kernel space vector comprises: The new kernel space vector is obtained by extracting features from the whitening results through fast independent principal component analysis.
5. An electronic device, characterized in that: The electronic device includes a memory, a processor, and a fault detection program based on an improved KECA stored in the memory and executable on the processor. When the fault detection program based on the improved KECA is executed by the processor, the steps of the fault detection method based on the improved KECA as described in any one of claims 1 to 4 are implemented.
6. A readable storage medium, characterized in that: The readable storage medium stores a fault detection program based on an improved KECA. When the fault detection program based on the improved KECA is executed by a processor, the steps of the fault detection method based on the improved KECA as claimed in any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Comprehensive anomaly extraction method based on blind source separation technology and apparatus thereof
CN103645517A
Mixed-norm multiple indefinite kernel classification method for complex data
CN106022382A