Uncertain industrial process robust fault diagnosis method based on interval distribution analysis
Through the embedded interval distribution analysis-radial basis function neural network algorithm, the problems of interval value data dimensionality reduction and structural information extraction in complex industrial processes are solved, and robust fault diagnosis of uncertain data is achieved, and the accuracy and robustness of fault identification are improved.
Patent Information
- Application Number
- CN202510247109.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-03
AI Technical Summary
The prior art is difficult to effectively reduce the dimension of interval value data and extract its internal structural information, resulting in a reduction in the accuracy of fault diagnosis of complex industrial processes.
An embedded interval distribution analysis-radial basis function neural network algorithm (IDA-RBFNN) is proposed. The uncertainty data is converted into interval value data through the kernel density estimation model. The structural information of the interval data is extracted using the complete information principal component analysis algorithm, and an interval radial basis function neural network is constructed for fault diagnosis.
It realizes robust fault diagnosis of uncertain data in complex industrial processes, improves the accuracy and robustness of fault identification, and maintains efficient fault detection capabilities in high uncertainty environments.
Smart Images

Figure CN120086705A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of complex industrial process fault diagnosis, and more specifically to real-time online fault diagnosis of complex industrial systems contaminated by uncertainty based on an interval uncertainty data-driven algorithm. Background Art
[0002] In the actual industrial production process, due to factors such as measurement noise interference, sensor aging, and harsh industrial production scenarios, the process data collected by sensors often contains interference from uncertain factors [1]. For example, in the hydrometallurgical process, underwater sensors are often covered by mud, resulting in inaccurate pressure measurement of the thickener [2]. Although these uncertainties may be fuzzy or unknown, they are usually bounded. By utilizing the bounded characteristics of uncertainties, interval-valued data can better reveal the hidden information compared to single-valued data. Therefore, for complex industrial processes contaminated by uncertainties, by converting single-valued process data into interval-valued data, process monitoring and fault diagnosis can be effectively performed [3].
[0003] The converted interval-valued data may reduce the accuracy of subsequent fault diagnosis due to still containing too much redundant information. Therefore, how to reduce the dimension of the converted interval-valued data and extract key information from it has become an urgent problem to be solved. Currently, methods for reducing the dimension of interval-valued data include vertices principal component analysis (VPCA), interval-valued functional principal component analysis (IFPCA), etc. The VPCA method was proposed by Cazes et al. [4]. By connecting all vertex matrices, the interval-valued data matrix is converted into a vertex coding matrix, and then the traditional PCA method is applied to reduce the dimension of the vertex coding matrix and calculate the principal components. Sun et al. [5] proposed the IFPCA method, which introduces a time-varying distance function containing midpoint and radius information to reduce information loss. However, these methods do not consider the internal distribution information of interval-valued data and fail to fully capture the structural information inside the interval. In addition, existing mainstream fault diagnosis methods are mainly based on traditional single-valued data diagnosis algorithms [6], making it difficult to effectively develop a fault diagnosis model for interval-valued data. Therefore, fully considering the uncertainty of process data and establishing an interval data fault diagnosis model based on uncertain data that can accurately capture the internal distribution information of the interval, so as to achieve accurate identification of faults generated in complex production processes, is an urgent problem to be solved, and research in related aspects has important theoretical significance and engineering value.
[0004] References:
[0005] [1] J.Yuan, S.Wang, F.L.Wang, and S.M.Zhang, Abnormal condition identification via OVR-IRBF-NN for the process industry with imprecise data and semantic information, Industrial and Engineering Chemistry Research, vol.59, no.11, pp.5072-5086, 2020.
[0006] [2] J.Yuan, S.Zhang, S.Wang, F.Wang, and L.Zhao, Process abnormity identification by fuzzy logic rules and expert estimated thresholds derived certainty factor, Chemometric Intell.Lab.Syst., vol.209, 2021, Art.no.104232.
[0007] [3] M.Mansouri, K.Dhibi, M.Hajji, K.Bouzara, H.Nounou, and M.Nounou, Interval-valued reduced RNN for fault detection and diagnosis for wind energy conversion systems, IEEE Sensors Journal, vol.22, no.13, pp.13581-13588, 2022.
[0008] [4] P.Cazes, A.Chouakria, E.Diday and Y.Schektman, Extension de l’analyse en composantes principales à des données de type intervalle, Revue de Statistique Appliquée, vol.45, no.3, pp.5-24, 1997.
[0009] [5]L.R.Sun, K.L.Wang, L.N.Xu, C.H.Zhang and T.Balezentis, A time-varying distance based interval-valued functional principal component analysis method - A case study of consumer price index, Information Sciences, vol.589, pp.94-116, 2022.
[0010] [6]L.Wen, X.Y.Li and L.Gao, A new two-level hierarchical diagnosis network based on convolutional neural network, IEEE Transactions on Instrumentation and Measurement, vol.69, no.2, pp.330-338, 2020. Summary of the Invention
[0011] The present invention proposes a robust fault diagnosis method for uncertain industrial processes based on interval distribution analysis, which is used to identify anomalies in the case of data uncertainty in a complex industrial environment. This fault diagnosis method is called the embedded interval distribution analysis - radial basis function neural network algorithm (IDA - RBFNN). First, on the premise of considering industrial costs and method effectiveness, extract the potential information embedded in uncertainty, establish an interval representation model of uncertain data based on the kernel density estimation model, convert the uncertain data collected from the complex industrial site into interval - valued data, comprehensively capture the internal structural characteristics of the data object, and effectively represent the inherent uncertainty in single - valued data. Subsequently, a data mining algorithm based on complete information principal component analysis (CIPCA) is constructed for interval data. Utilize the inherent internal distribution characteristics of interval - valued data to construct a covariance matrix through inner product and squared norm, thereby extracting the structural information in the key interval fault features, realizing multivariate variable dimension reduction, and obtaining interval fault feature data. Finally, an interval radial basis function neural network (IRBFNN) is constructed based on the subtractive clustering algorithm. By processing the interval upper and lower bound matrices, cluster the interval fault feature data, and deeply analyze the hidden fault information under uncertainty, thereby improving the accuracy of fault diagnosis in industrial processes contaminated by uncertainty. The technical solutions are as follows:
[0012] A robust fault diagnosis method for uncertain industrial processes based on interval distribution analysis, comprising the following steps:
[0013] Step 1, establish an interval representation model of uncertain data based on the kernel density estimation model: For the process data collected by sensors in the industrial process and disturbed by uncertainty, use the kernel density estimation model to fit the deviation distribution between the sensor measurement value and the true value of the process variable, describe the process data disturbed by uncertainty using interval - type symbolic data, and uniformly convert the process data collected in the industrial process and disturbed by uncertainty into an interval form;
[0014] Step 2, extraction of uncertain process characteristics based on the complete information principal component analysis algorithm: Use the uncertainty data - driven algorithm based on complete information principal component analysis to mine the potential structural information characteristics of the uncertain process data represented in interval form, so as to realize the extraction of uncertain fault feature information and multivariate variable dimension reduction;
[0015] Step 3, construct a fault diagnosis model based on the interval radial basis function neural network algorithm: Based on the interval fault feature information extracted by the complete information principal component analysis algorithm constructed in the second part, construct a diagnosis model applicable to interval-valued fault data, and use the subtractive clustering algorithm to complete the offline modeling of the diagnosis model;
[0016] Step 4, real-time online diagnosis of industrial process faults: During the operation of the industrial process, process information is collected online, and the uncertainty process data is uniformly converted into an interval form by using the established interval representation model of uncertainty data based on the kernel density estimation model; Subsequently, use the complete information principal component analysis algorithm to extract the potential structure feature information of the online interval process data, and calculate the online fault feature information; Finally, input the online fault feature information into the interval radial basis function neural network fault diagnosis model to determine which specific fault has occurred in the system.
[0017] Further, the method of Step 1 is as follows:
[0018] Record the process data collected under the fault condition as where m represents the number of samples, n represents the number of process variables, and x j =[x 1j , x 2j ,..., x mj T is the measured value of the j-th process variable, and its corresponding true value is expressed as Then the corresponding measurement deviation ε j is:
[0019]
[0020] Based on the kernel density estimation model, the probability density distribution function j of the measurement deviation ε is:
[0021]
[0022] where h represents the kernel bandwidth, and K(·) selects the Gaussian kernel;
[0023] When the significance level is taken as α, calculate the upper limit j and the lower limit j of the measurement deviation ε ε , and further quantify the error range to evaluate the uncertainty caused by the measurement error:
[0024]
[0025] The i-th sample of the j-th process variable is described as interval-valued data as follows:
[0026]
[0027] Based on the established kernel density estimation model, the process data with uncertainty is uniformly represented as an interval-valued data matrix:
[0028]
[0029] where, [x j represents the interval observation values of all samples of the j-th process variable.
[0030] Furthermore, the method of step 2 is as follows:
[0031] (1) Interval data standardization:
[0032]
[0033] where, E([x j ) represents the mean value of the j-th process variable;
[0034] (2) Extracting eigenvectors based on the interval covariance matrix:
[0035]
[0036] where, <[x ij ,[x ig > represents the inner product between the i-th sample points corresponding to the j-th and g-th process variables;
[0037] Perform eigenvalue decomposition on the covariance matrix of the interval data according to the following formula:
[0038] Σ CI =PΛP T (8)
[0039] Determine the parameter l according to the cumulative variance percentage CPV criterion, and obtain the reduced-dimensional eigenvector P = [p 1 ,p 2 ,...,p l , (l ≤ n) and the corresponding eigenvalues λ 1 ≥λ 2 ≥...≥λ l ;
[0040] (3) Extraction of fault feature data: Extract the key interval fault feature data [X d = ([t 1 ,[t 2,...,[t l ):
[0041]
[0042] where \(h = 1, 2,\cdots, l\), \(p\) kh is the \(k\)th element in the \(h\)th eigenvector \(p\) h .
[0043] Furthermore, the method of step 3 is as follows:
[0044] Based on the interval radial basis function neural network, the interval fault feature data \([X\) d = ([t 1 , [t 2 , \cdots, [t l ) is classified for faults. The density index matrix \(D = [d\) 1 , d\) 2 , \cdots, d\) m between sample points is calculated as follows:
[0045]
[0046] where \(\|[s\) i - [s\) j \| represents the Euclidean distance between the \(i\)th sample point and the \(j\)th sample point in the interval fault feature data, and \(\gamma\) a represents a related constant coefficient;
[0047] Assume that the density index \(d\) q of the \(q\)th sample point is the maximum value of the entire density index matrix. Then the first interval clustering center vector \([v\) 1 and the corresponding density index \(d\) v can be expressed as:
[0048]
[0049] The density index matrix is iteratively updated to obtain a new interval clustering center vector and density index. The corresponding iterative process is as follows:
[0050]
[0051] where \(s\) pr represents the interval vector of the interval clustering center vector before iteration, \(r\) represents the \(r\)th iteration, and \(\gamma\) b is a constant coefficient related to \(\gamma\) a ;
[0052] New clustering centers are continuously obtained until the following termination condition is met:
[0053]
[0054] Among them, ρ represents the iteration termination constant coefficient, which determines whether to continue the iteration of the density index matrix;
[0055] After the iteration termination, all interval clustering center vectors [V] = ([v 1 ,[v 2 ,...,[v r ) and the corresponding maximum density index matrix are obtained r interval clustering center vectors represent r nodes in the hidden layer of the neural network;
[0056] By setting the maximum number of iterations c and the learning rate θ, and using the gradient descent algorithm to train the linear weight matrix ω from the hidden layer to the output layer using the error E between the predicted value and the actual value:
[0057]
[0058] Among them, ω i represents the i-th iteration value of the weight matrix ω;
[0059] The fault feature data outputs the label of the predicted fault classification through the calculation of the interval radial basis neural network, completing the fault diagnosis of the uncertain industrial process.
[0060] The method for interval representation of uncertain data proposed by the present invention introduces a kernel density estimation model to fit the probability density function of measurement errors. On the basis of in-depth analysis of the sources and characteristics of uncertainties, a scientific and reliable method is realized to accurately convert the uncertain information contained in process data into the form of interval-type symbolic data. At the same time, compared with the existing complex industrial fault diagnosis algorithms, the embedded interval distribution analysis-radial basis function neural network algorithm designed by the present invention fully considers the distribution information inside the interval data, accurately extracts and clusters the key interval fault feature data, and constructs a robust fault diagnosis model under the interference of uncertain information. Brief Description of the Drawings
[0061] Figure 1 Schematic Diagram of the Robust Fault Diagnosis Algorithm for Uncertain Industrial Processes Based on Interval Distribution Analysis
[0062] Figure 2 Process Flow Chart of the TE System
[0063] Figure 3 T-SNE Visualization Diagram of Different Interval Dimensionality Reduction Algorithms under Interval Test Data
[0064] Figure 4 Performance of Each Fault of Different Fault Diagnosis Algorithms under an Uncertainty Level of 1% (%)
[0065] Figure 5 Fault diagnosis diagrams for each fault of different fault diagnosis algorithms under an uncertainty level of 1%
[0066] Figure 6 Performance of different diagnostic algorithms for all fault types under four uncertainty levels (%)
[0067] Figure 7 Performance comparison diagrams of different diagnostic algorithms for all fault types under four uncertainty levels Detailed implementation manners
[0068] The present invention proposes a robust fault diagnosis method for an uncertain industrial process based on interval distribution analysis, which is used for anomaly identification under the condition of data uncertainty in a complex industrial environment. Specifically, as Figure 1 shown, this diagnosis method constructs a robust fault diagnosis model for processing uncertainty interference by processing interval process data. The entire fault diagnosis system mainly includes the following four parts: interval representation processing of uncertain data, uncertainty process feature extraction based on the complete information principal component analysis algorithm, construction of a fault diagnosis model based on the interval radial basis function neural network algorithm, and real-time online diagnosis of industrial process faults
[0069] First, the basic technical solution of the present invention is introduced
[0070] The robust fault diagnosis method for an uncertain industrial process based on interval distribution analysis of the present invention includes the following steps
[0071] Step 1, interval representation processing of uncertain data: For the process data affected by uncertainty interference collected by sensors in the industrial process, use the kernel density estimation model to fit the deviation distribution between the sensor measurement value and the true value of the process variable, and use interval-type symbolic data to describe the process data affected by uncertainty interference, so as to uniformly convert the process data affected by uncertainty interference collected in the industrial process into an interval form, improving the robustness to uncertainty interference. Specifically as follows
[0072] Record the process data collected under the fault condition as where m represents the number of samples, n represents the number of process variables, and x j =[x 1j , x 2j ,..., x mj T is the measured value of the jth process variable, and its corresponding true value is expressed as Then the corresponding measurement deviation ε j is
[0073]
[0074] Based on the kernel density estimation model, measure the deviation ε j , j = 1, 2, ..., n, of the probability density distribution function is:
[0075]
[0076] where h represents the kernel bandwidth and K(·) selects the Gaussian kernel;
[0077] When the significance level is taken as α, it can be calculated that the upper limit of the measurement deviation ε j and the lower limit ε and further quantify the error range to evaluate the uncertainty caused by the measurement error: j The i-th sample point of the j-th process variable can be described as interval-valued data as shown below:
[0078]
[0079] Based on the established kernel density estimation model, the process data with uncertainty is uniformly represented as an interval-valued data matrix:
[0080]
[0081] where [x
[0082]
[0083] represents the interval observation values of all samples of the j-th process variable. j represents the interval observation values of all samples of the j-th process variable.
[0084] Step 2, Uncertainty process feature extraction based on the complete information principal component analysis algorithm: Use the uncertainty data-driven algorithm based on the complete information principal component analysis to mine the potential structural information features of the uncertainty process data represented in interval form, so as to realize the extraction of uncertainty fault feature information and the dimensionality reduction of multivariate variables. Specifically as follows:
[0085] (1) Standardize the interval data; Standardize the obtained interval data and keep the corresponding representation symbols of the process data unchanged before and after standardization:
[0086]
[0087] where E([x j ) represents the mean value of the j-th process variable;
[0088] (2) Extract feature vectors based on the interval covariance matrix; to extract potential structural information, calculate the covariance matrix for the interval data after standardization according to the distribution within the interval:
[0089]
[0090] where <[x ij ,[x ig > represents the inner product between the i-th sample points corresponding to the j-th and g-th process variables;
[0091] Subsequently, perform eigenvalue decomposition on the covariance matrix of the interval data according to the following formula:
[0092] Σ CI =PΛP T (8)
[0093] Finally, determine the parameter l according to the cumulative percentage of variance CPV criterion, and obtain the reduced-dimensional feature vector P = [p 1 ,p 2 ,...,p l , (l ≤ n) and the corresponding eigenvalues λ 1 ≥λ 2 ≥...≥λ l ;
[0094] (3) Extraction of fault feature data: Extract the key fault feature data [X d = ([t 1 ,[t 2 ,...,[t l ) from the above feature vector matrix based on Moore's law of linear combination:
[0095]
[0096] where h = 1, 2,... l, and p kh is the k-th element in the h-th feature vector p h .
[0097] Step 3, construct a fault diagnosis model based on the interval radial basis function neural network algorithm: Based on the interval fault feature information extracted by the complete information principal component analysis algorithm constructed in the second part, construct a diagnosis model suitable for interval-valued fault data, and use the subtractive clustering algorithm to complete the offline modeling of the diagnosis model.
[0098] Specifically as follows:
[0099] Based on the interval radial basis function neural network, use the subtractive clustering algorithm for the interval fault feature data [X d = ([t 1 ,[t2 ,...,[t l ), the density index matrix D = [d 1 , d 2 ,..., d m between sample points is calculated as follows:
[0100]
[0101] where, ||[s i - [s j || represents the Euclidean distance between the i-th sample point and the j-th sample point in the interval fault feature data, and γ a represents a relevant constant coefficient;
[0102] Assume that the density index d q of the q-th sample point is the maximum value of the entire density index matrix. Then, the first interval clustering center vector [v 1 and the corresponding density index d v can be expressed as follows:
[0103]
[0104] Iteratively update the density index matrix to obtain a new interval clustering center vector and density index. The corresponding iterative process is as follows:
[0105]
[0106] where, s pr represents the interval vector of the interval clustering center vector before iteration, r represents the r-th iteration, and γ b is a constant coefficient related to γ a ;
[0107] Continuously obtain new clustering centers until the following termination condition is met:
[0108]
[0109] where, ρ represents the iterative termination constant coefficient, which determines whether the density index matrix continues to be iterated;
[0110] After the iteration terminates, all interval clustering center vectors [V] = ([v 1 , [v 2 ,..., [v r ) and the corresponding maximum density index matrix The r interval clustering center vectors represent r nodes in the hidden layer of the neural network.
[0111] By setting the maximum number of iterations \(c\) and the learning rate \(\theta\), and using the gradient descent algorithm to train the linear weight matrix \(\omega\) from the hidden layer to the output layer by using the error \(E\) between the predicted value and the actual value:
[0112]
[0113] where \(\omega\) i represents the \(i\)-th iteration value of the weight matrix \(\omega\).
[0114] Finally, the fault feature data outputs the label of the predicted fault classification through the calculation of the interval radial basis neural network, completing the fault diagnosis of the uncertain industrial process.
[0115] Step 4, real-time online diagnosis of industrial process faults: During the operation of the industrial process, process information is collected online, and the established uncertainty data intervalization characterization model based on the kernel density estimation model is used to uniformly convert the uncertain process data into an interval form; subsequently, the potential structure feature information of the online interval process data is extracted by using the complete information principal component analysis algorithm, and the online fault feature information is calculated; finally, the online fault feature information is input into the interval radial basis function neural network fault diagnosis model to determine which specific fault has occurred in the system.
[0116] The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of this application. The application object comes from a chemical process simulation model developed by Tennessee Eastman (TE) Company. The specific process is as Figure 2 shown, with a total of 52 process variables and 21 types of fault types. The first four types of faults are adopted in this embodiment. Each group of faults contains 480 training samples and 200 test samples, and 4 levels of uncertainty \(\eta\) are designed.
[0117] Step 1, intervalization characterization processing of uncertain data: For the process data affected by uncertainty collected by sensors in the actual production process, the kernel density estimation model is used to fit the deviation distribution between the sensor measurement values and the true values of different process variables, and interval type symbolic data is used to describe the process data affected by uncertainty, so as to uniformly convert the process data affected by uncertainty collected in the industrial process into an interval form, improving the robustness to uncertainty interference.
[0118] The training samples collected under the four types of fault conditions are concatenated and denoted as where 1920 represents the total number of samples, 52 represents the number of process variables, and the following uncertainty interferences are added:
[0119]
[0120] Subsequently, x was obtained j = [x 1j , x 2j ,..., x 1920j T is the measured value of the j-th process variable, where j = 1, 2,..., 52, and its corresponding true value is denoted as Then the corresponding measurement deviation ε j is:
[0121]
[0122] Based on the kernel density estimation model, the probability density distribution function j of the measurement deviation ε is:
[0123]
[0124] where h represents the kernel bandwidth, which depends on the sample data of each process variable, and K(·) selects the Gaussian kernel;
[0125] When the significance level is taken as 0.95, the upper limit j and the lower limit of the measurement deviation ε ε j can be calculated, and then the error range can be quantified to evaluate the uncertainty caused by the measurement error:
[0126]
[0127] The i-th sample point of the j-th process variable can be described as interval-valued data as shown below:
[0128]
[0129] Based on the established kernel density estimation model, the process data containing uncertainty is uniformly represented as an interval-valued data matrix:
[0130]
[0131] where [x j represents the interval observations of all samples of the j-th process variable.
[0132] Step 2, Uncertainty process feature extraction based on the complete information principal component analysis algorithm: Use the uncertainty data-driven algorithm based on the complete information principal component analysis to mine the potential structural information features of the uncertainty process data represented in interval form, so as to realize the extraction of uncertainty fault feature information and multivariate variable dimension reduction. Taking the uncertainty level of 1% as an example, the specific operations are as follows:
[0133] (1) Interval data standardization processing; perform standardization processing on the obtained interval data and keep the corresponding representation symbols of the process data unchanged before and after standardization:
[0134]
[0135] where E([x j ) represents the mean value of the j-th, j = 1, 2,..., 52 process variables;
[0136] (2) Extracting eigenvectors based on the interval covariance matrix; in order to extract potential structure information, calculate the covariance matrix for the standardized interval data according to the within-interval distribution, and the result is as follows:
[0137]
[0138] Subsequently, select the cumulative variance contribution rate of 0.999 to perform eigen-decomposition on the covariance matrix of the interval data to obtain the eigenvector matrix P and the corresponding eigenvalue matrix Λ:
[0139]
[0140] (3) Extraction of fault feature data: Based on Moore's linear combination law, extract the key fault feature data [X d from the above eigenvector matrix P as follows:
[0141]
[0142] Step 3, constructing a fault diagnosis model based on the interval radial basis function neural network algorithm: Based on the extracted key interval fault feature data, construct a diagnosis model suitable for interval-valued fault data, and use the subtractive clustering algorithm to complete the offline modeling of the diagnosis model:
[0143] The termination constant coefficient ρ in the interval subtractive clustering algorithm of the hidden layer is 0.3, and set the learning rate of the weight matrix for training the hidden layer to the output layer as θ = 0.001, and obtain the hidden layer center vector matrix [V] and the weight matrix W as follows:
[0144]
[0145] Step 4, real-time online diagnosis of industrial process faults: During the operation of the industrial process, collect process information online, and use the established uncertainty data intervalization characterization model based on the kernel density estimation model to uniformly convert the uncertain process data into interval form; subsequently, use the complete information principal component analysis algorithm to extract the potential structure feature information of the online interval process data and calculate the online fault feature information; finally, input the online fault feature information into the interval radial basis function neural network fault diagnosis model to determine which specific fault has occurred in the system.
[0146] Taking the uncertainty level of 1% as an example, the test set data is substituted into the interval representation model in Step 1 to be converted into interval data Subsequently, the interval data is substituted into the data dimensionality reduction model in Step 2, and the key fault feature data of the test set is calculated using the known feature vector matrix P and Moore's law of linear combination Finally, the fault feature data is input into the established interval radial basis neural network model, and the fault label information corresponding to each fault sample is obtained, completing the entire fault diagnosis process.
[0147] To verify the superiority of CIPCA in feature extraction of interval data, comparisons are made with the original data and CPCA respectively, as Figure 3 shown. The circle points represent the upper limit data of the interval, and the cross points represent the lower limit data of the interval. It can be seen that the data processed by CIPCA can be better classified. In addition, to verify the superiority of IDA-RBFNN, comparisons are made with the fault diagnosis performances of BPNN, RBFNN, IRBFNN, and CPCA-IRBFNN respectively, as Figures 4 to 7 shown. The simulation results show that the recognition accuracy of the IDA-RBFNN method for various faults is not less than 90%, and the overall accuracy reaches 97.5%, and its performance is better than other comparison methods. Further, to explore the influence of the interference of uncertainty on the final fault diagnosis, the research is carried out by gradually increasing the uncertainty level. Uncertainty levels of 1%, 10%, 20%, and 30% are successively introduced into different fault diagnosis models. The experimental data show that as the uncertainty level increases, the fault diagnosis accuracy of all models shows a downward trend. When the average decline rate of the four comparison algorithms reaches 16.85%, the decline rate of the algorithm of the present invention is only 7.75%, indicating that the IDA-RBFNN algorithm has better robustness. It should be noted that at different uncertainty levels, the overall accuracy of IRBFNN is always higher than that of RBFNN. Especially when the uncertainty level is increased to 30%, that is, when the interference is the strongest, the accuracy of IRBFNN is significantly increased by 19.75% compared with RBFNN, fully verifying the necessity of data intervalization processing.
[0148] The applicant of the present invention has made a detailed description and explanation of the implementation examples of the present invention in combination with the accompanying drawings of the specification. However, those skilled in the art should understand that the above implementation examples are only the preferred implementation schemes of the present invention, and the detailed description is only to help readers better understand the spirit of the present invention, rather than a limitation on the protection scope of the present invention. On the contrary, any improvement or modification based on the spirit of the present invention should fall within the protection scope of the present invention.
Claims
1. A robust fault diagnosis method for uncertain industrial processes based on interval distribution analysis, comprising the following steps: Step 1: Establish an interval characterization model for uncertainty data based on the kernel density estimation model: For the process data interfered by uncertainty collected by sensors in the industrial process, the kernel density estimation model is used to fit the deviation distribution between the sensor measurement value and the true value of the process variable, and the interval symbolic data is used to describe the process data interfered by uncertainty, and the process data interfered by uncertainty collected in the industrial process are uniformly converted into interval form; Step 2, uncertainty process feature extraction based on complete information principal component analysis algorithm: The uncertainty data driven algorithm based on complete information principal component analysis is used to mine the potential structural information features of the uncertainty process data represented in interval form, so as to realize the extraction of uncertainty fault feature information and multivariate variable dimensionality reduction; Step 3, construct a fault diagnosis model based on interval radial basis function neural network algorithm: Based on the interval fault feature information extracted by the complete information principal component analysis algorithm constructed in the second part, a diagnosis model suitable for interval value fault data is constructed, and the off-line modeling of the diagnosis model is completed by using the clustering reduction algorithm; Step 4, real-time online diagnosis of industrial process faults: When the industrial process is running, the process information is collected online, and the uncertainty data interval representation model based on the kernel density estimation model is used to uniformly convert the uncertainty process data into interval form; then, the potential structural feature information of the online interval process data is extracted by the principal component analysis algorithm based on complete information, and the online fault feature information is calculated; finally, the online fault feature information is input into the interval radial basis function neural network fault diagnosis model to determine what specific fault has occurred in the system.
2. The method for robust fault diagnosis of uncertain industrial processes based on interval distribution analysis according to claim 1 is characterized in that: The method for step 1 is as follows: The process data collected under fault conditions is recorded as Where m is the number of samples, n is the number of process variables, and x is j =[x 1j ,x 2j ,…,x mj ] T is the measured value of the jth process variable, and its corresponding true value is expressed as The corresponding measurement deviation ε j for: Based on the kernel density estimation model, the measurement deviation ε j ,j=1,2,...,n,probability density distribution function for: Where h represents the kernel bandwidth, and K(·) selects the Gaussian kernel; When the significance level is α, the measurement deviation ε is calculated j Upper limit and lower limit ε j , and then quantify the error range to assess the uncertainty caused by measurement error: The i-th sample of the j-th process variable is described as interval-valued data as shown below: Based on the established kernel density estimation model, the process data with uncertainty is uniformly represented as an interval value data matrix: Among them, [x j ] represents the interval observation value of all samples of the jth process variable.
3. The method for robust fault diagnosis of uncertain industrial processes based on interval distribution analysis according to claim 2 is characterized in that: The method for step 2 is as follows: (1) Standardization of interval data: Among them, E([x j ]) represents the mean value of the jth process variable; (2) Extract eigenvectors based on interval covariance matrix: Among them, <[x ij ],[x ig ]> represents the inner product between the i-th sample points corresponding to the j-th and g-th process variables; The covariance matrix of interval data is characteristically decomposed according to the following formula: S CI =PΛP T (8) According to the cumulative variance percentage CPV criterion, the parameter l is determined and the eigenvector P after dimensionality reduction is obtained. l ], (l≤n) and the corresponding eigenvalues λ1≥λ2≥...≥λ l ; (3) Extraction of fault feature data: Based on Moore's law of linear combinations, the key interval fault feature data [X d ]=([t1],[t2],...,[t l ]): Where h = 1, 2, ... l, p kh is the hth eigenvector p h The kth element in .
4. The method for robust fault diagnosis of uncertain industrial processes based on interval distribution analysis according to claim 3 is characterized in that: The method for step 3 is as follows: Based on interval radial basis function neural network, the subtractive clustering algorithm is used to analyze the interval fault feature data[X d ]=([t1],[t2],...,[t l ]) to classify faults, the density index matrix D between sample points = [d1, d2, ..., d m ] is calculated as follows: Among them, ||[s i ]-[s j ]|| represents the Euclidean distance between the i-th sample point and the j-th sample point in the interval fault feature data, γ a represents a correlation constant coefficient; Assume that the density index d of the qth sample point is q is the maximum value of the entire density index matrix, then the first interval cluster center vector [v1] and the corresponding density index d v They can be expressed as: The density index matrix is iteratively updated to obtain new interval cluster center vectors and density indexes. The corresponding iterative process is as follows: Among them, s pr represents the interval cluster center vector interval vector before iteration, r represents the rth iteration, γ b For a A constant coefficient of correlation; New cluster centers are continuously obtained until the following termination conditions are met: Among them, ρ represents the iteration termination constant coefficient, which determines whether the density indicator matrix continues to iterate; After the iteration is terminated, all the interval cluster center vectors [V] = ([v1], [v2], ..., [v r ]) and the corresponding maximum density indicator matrix The r interval cluster center vectors represent the r nodes in the hidden layer of the neural network; By setting the maximum number of iterations c and the learning rate θ, and using the gradient descent algorithm to train the linear weight matrix ω from the hidden layer to the output layer using the error E between the predicted value and the actual value: Among them, ω i Represents the i-th iteration value of the weight matrix ω; The fault feature data is calculated through the interval radial basis function neural network to output the predicted fault classification label, thus completing the diagnosis of uncertain industrial process faults.