Industrial fault detection method based on joint sparse low-rank double dictionary learning

By employing a joint sparse low-rank dual-dictionary learning method, the adaptability problem of traditional fault detection methods in complex nonlinear systems is solved. Through collaborative learning of two sub-dictionary models, global and local structural information in industrial processes is extracted, thereby improving the accuracy and robustness of fault detection.

CN121456575APending Publication Date: 2026-02-03BEIJING GUODIAN ZHISHEN CONTROL TONGDY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511310741.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Traditional fault detection methods are poorly adapted to complex nonlinear systems and cannot meet the high precision, real-time and robustness requirements of modern industrial systems. Using sparse or low-rank representation alone is difficult to take into account both local details and global structure, which limits its application in complex industrial scenarios.

Method used

We adopt a dual-dictionary learning method based on joint sparse low-rank representation. By co-learning two sub-dictionaries and combining sparse and low-rank representations, we construct a dual-dictionary learning model. We use an alternating optimization mechanism to solve the objective function, reduce noise interference, and extract global and local structural information from the data.

Benefits of technology

It significantly improves the industrial fault detection rate, effectively detects nonlinearity, noise interference and high-dimensional data in complex industrial processes, and enhances the fault detection performance of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121456575A_ABST
    Figure CN121456575A_ABST
Patent Text Reader

Abstract

The invention discloses an industrial fault detection method based on joint sparse low-rank double-dictionary learning, relates to the technical field of engineering fault detection, and is technically characterized in that sparse representation and low-rank representation are combined, single-dictionary learning is expanded, and the industrial fault detection method based on double-dictionary learning is constructed; according to the method, the problems of nonlinearity, noise interference, high-dimensional data and the like in a complex industrial process are solved, and the fault detection rate of the system is remarkably improved; in a high-dimensional complex industrial process including noise pollution and a multivariable coupling relationship, the scheme can effectively detect system faults.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of engineering fault detection technology, specifically to an industrial fault detection method based on joint sparse low-rank dual dictionary learning. Background Technology

[0002] With the increasing complexity and intelligence of industrial systems in the process industry, equipment operation status monitoring and fault detection have become crucial for ensuring system safety and improving production efficiency. Traditional fault detection methods mainly rely on expert knowledge and empirical rules, which have the disadvantage of poor adaptability to complex nonlinear systems and difficulty in meeting the high precision, real-time performance, and robustness requirements of modern industrial systems. In recent years, data-driven methods have been widely used in industrial fault detection, especially machine learning-based models, which have shown significant advantages in tasks such as anomaly detection and fault diagnosis.

[0003] In real-world industrial environments, information redundancy and noise pollution exist. Sparse representation and low-rank representation theories have been proposed and widely used to address these issues. The former eliminates redundant information through dimensionality reduction and reconstruction of signals based on their sparsity, while the latter mines the global structure of data and suppresses the impact of noise. Dictionary learning, by constructing a dictionary adapted to the data, can extract hidden features and avoids the dependence of traditional methods on the linear distribution and noise characteristics of the data. However, using sparse or low-rank representation alone makes it difficult to simultaneously consider local details and global structure, limiting its application in complex industrial scenarios.

[0004] Therefore, the present invention aims to provide an industrial fault detection method based on joint sparse low-rank dual dictionary learning to solve the above problems. Summary of the Invention

[0005] The purpose of this invention is to solve the above-mentioned problems and provide an industrial fault detection method based on joint sparse low-rank dual dictionary learning. By combining sparse representation and low-rank representation and extending single dictionary learning, it solves problems such as nonlinearity, noise interference and high-dimensional data in complex industrial processes, and significantly improves the fault detection rate of the system.

[0006] To achieve the above objectives, the technical solution of the present invention is as follows:

[0007] This invention provides an industrial fault detection method based on joint sparse low-rank dual-dictionary learning. 1. Establishment of the model objective function.

[0008] In industrial process monitoring, the raw data X∈R N×d Containing N samples with dimensional observation variables, in a traditional single-dictionary learning model, the original data can be decomposed into the product of a dictionary D and an encoding matrix Y:

[0009]

[0010] Where ||Y||1 is a sparse constraint term to ensure the discriminative nature of the representation learning. To ensure that the learned dictionary can flexibly encode different information, this invention simultaneously learns two sub-dictionaries D1 and D2, respectively learning the encoding matrices Y1 and Y2 corresponding to the two sub-dictionaries, thereby reflecting and encoding different structural information. Through the collaborative learning of the two sub-dictionaries, richer process measurement key information can be included in the encoding matrix. Since the two sub-dictionaries are learned collaboratively, the content learned by the two sub-dictionaries is interrelated; therefore, the resulting dictionaries have a good cooperative relationship.

[0011] Data from industrial processes typically exhibits characteristics such as high dimensionality, complexity, and high noise levels, while containing structural information that is interconnected both globally and locally. Different failure modes are usually distributed across diverse feature spaces. Therefore, dual-dictionary learning models, by combining low-rank and sparse constraints, can effectively extract global geometric structure (low-rank characteristics) and local geometric structure (sparse characteristics) from the data, reducing noise interference and uncovering key information.

[0012] For low-rank properties, the kernel norm is generally used to constrain the matrix; for sparsity, the norm is usually used as a constraint. Based on the classic dictionary learning model (1) and inspired by dual-dictionary learning, a dual-dictionary learning model can be constructed:

[0013]

[0014] Where D1 and Y1 represent the low-rank dictionary learning part, D2 and Y2 represent the sparse representation dictionary learning part, and the nuclear norm is ||·|| * Constrain low-rank structures and l1 norm constraints on sparse structures.

[0015] To improve the discriminative power of the dictionary, this invention represents the dictionary as the product of the original data and the matrix composed of the dictionary, expressed as:

[0016]

[0017] However, for low-rank constraints, constraining only the dictionary can ensure that the information in the low-rank part is fully preserved, therefore a||XW1Y1|| * The term can be ignored, and for the l1 norm constraints of Y1 and Y2, l is used. 2,1 Norm constraints are replaced to enhance feature selection capabilities. The final objective function of the joint sparse low-rank dual-dictionary learning model can be expressed as:

[0018]

[0019] Where a, b, c, and d are the balance parameters of the corresponding terms.

[0020] 2. Objective function optimization

[0021] For the objective function of the established model, this is a non-convex problem with respect to W1, W2, Y1, and Y2; however, it is a convex problem for any one of these variables. In this invention, an alternating optimization mechanism is used to solve the objective function (4). The idea of ​​the alternating iterative optimization mechanism is to fix the remaining variables and solve one variable. After obtaining the solution, the variable being solved is fixed, and the optimization problem of the previously fixed variable is solved until a relatively stable local optimum is obtained. In order to solve the objective function, Lagrange multipliers are introduced to constrain the variables W1, W2, Y1, and Y2, and constraint terms are added. Construct the Lagrange function:

[0022]

[0023] Where η and μ are penalty parameters, for ||A|| 2,1 Introducing a diagonal matrix U ii = 1 / 2||A||2, and initialize it as an identity matrix. Then ||A|| 2,1 Transform into matrix trace Tr(A) T The form is UA). For ||A||1, introduce two all-1 vectors i. c and i d Transform ||A||1 into Tr(i) c T A T i d For the nuclear norm ||XW1|| * Its physical meaning can be interpreted as the sum of the singular values ​​of the matrix, therefore we have ||XW1|| * =Tr(Σ), and define XW1 = MΣN T , and are the left singular matrix and the right singular matrix, respectively, and both are orthogonal matrices.

[0024] Referring to the above method and the transformation relationship between norm and matrix trace, equation (5) can be transformed into the form of matrix trace:

[0025]

[0026] U1 and U2 represent Y1 and Y2 respectively as l 2,1 The auxiliary diagonal matrix introduced by arithmetic operations is derived from the following relationship:

[0027]

[0028] Then we begin updating each variable. First, we fix W2, Y1, and Y2, update matrix W1, and calculate the partial derivatives of equation (6).

[0029]

[0030] Use KKT optimization conditions to let

[0031]

[0032] achievable

[0033]

[0034] Further, the update and iteration rules of W1 are obtained.

[0035]

[0036] Similarly, the update rules for W2, Y1, and Y2 can be obtained. The partial derivatives of the objective function with respect to variables W2, Y1, and Y2 are calculated as follows:

[0037]

[0038] Use KKT optimization conditions to let Further, the update rules for W2, Y1, and Y2 are obtained.

[0039]

[0040] 3. Convergence conditions

[0041] The maximum number of iterations is set to 500, and the stopping criterion is defined as follows:

[0042]

[0043] Where D1 = XW1, D2 = XW2, when RelErr is 10 -3 Stop updating the optimization process when necessary.

[0044] 4. Fault detection applications and criteria

[0045] After learning and obtaining the dictionary matrices D1 and D2, the test data X is... new And by reconstructing the coefficient matrices Y1 and Y2, we obtain and

[0046]

[0047] T 2 Statistics measure the change of a sample vector in the principal space.

[0048]

[0049] To detect faults, if the test statistic exceeds the control limit, a fault is considered to have occurred; otherwise, no fault has occurred. Therefore, the test logic can be defined as follows:

[0050]

[0051] in, The control limit, representing a confidence level of α, is typically calculated using KDE (kernel density estimation).

[0052] To measure the performance of fault detection, a widely used metric is the alarm rate (FDR). The FDR is defined as follows:

[0053]

[0054] Where f = 0 means that no failure occurred in the test sample set, and f ≠ 0 means that a failure occurred in the test sample set.

[0055] Compared with existing technologies, the beneficial effects of this solution are:

[0056] This invention provides an industrial fault detection method based on joint sparse and low-rank dual dictionary learning (LSDDL). This method extends single dictionary learning and combines sparse representation and low-rank representation, providing an effective solution to problems such as nonlinearity, noise interference, and high-dimensional data in complex industrial processes, and can significantly improve the fault detection rate of the system. In high-dimensional complex industrial processes containing noise pollution and multivariate coupling relationships, this scheme can effectively detect system faults. Attached Figure Description

[0057] Figure 1 This is a flowchart illustrating fault detection using the LSDDL method in an embodiment of the present invention;

[0058] Figure 2 These are the statistical changes of the six selected faults in the embodiments of the present invention. Detailed Implementation

[0059] To enable those skilled in the art to better understand the present invention, the technical solution of the present invention will be described in further detail below with reference to the embodiments and accompanying drawings. Obviously, the described embodiments are merely some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort should fall within the scope of protection of the present invention.

[0060] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the embodiments.

[0061] Example 1:

[0062] 1. Establishment of the model objective function

[0063] In industrial process monitoring, the raw data X∈R N×d Containing N samples with dimensional observation variables, in a traditional single-dictionary learning model, the original data can be decomposed into the product of a dictionary D and an encoding matrix Y:

[0064]

[0065] Where ||Y||1 is a sparse constraint term to ensure the discriminative nature of the representation learning. To ensure that the learned dictionary can flexibly encode different information, this invention simultaneously learns two sub-dictionaries D1 and D2, respectively learning the encoding matrices Y1 and Y2 corresponding to the two sub-dictionaries, thereby reflecting and encoding different structural information. Through the collaborative learning of the two sub-dictionaries, richer process measurement key information can be included in the encoding matrix. Since the two sub-dictionaries are learned collaboratively, the content learned by the two sub-dictionaries is interrelated; therefore, the resulting dictionaries have a good cooperative relationship.

[0066] Data from industrial processes typically exhibits characteristics such as high dimensionality, complexity, and high noise levels, while containing structural information that is interconnected both globally and locally. Different failure modes are usually distributed across diverse feature spaces. Therefore, dual-dictionary learning models, by combining low-rank and sparse constraints, can effectively extract global geometric structure (low-rank characteristics) and local geometric structure (sparse characteristics) from the data, reducing noise interference and uncovering key information.

[0067] For low-rank properties, the kernel norm is generally used to constrain the matrix; for sparsity, the norm is usually used as a constraint. Based on the classic dictionary learning model (1) and inspired by dual-dictionary learning, a dual-dictionary learning model can be constructed:

[0068]

[0069] Where D1 and Y1 represent the low-rank dictionary learning part, D2 and Y2 represent the sparse representation dictionary learning part, and the nuclear norm is ||·|| * Constrain low-rank structures and l1 norm constraints on sparse structures.

[0070] To improve the discriminative power of the dictionary, this invention represents the dictionary as the product of the original data and the matrix composed of the dictionary, expressed as:

[0071]

[0072] However, for low-rank constraints, constraining only the dictionary can ensure that the information in the low-rank part is fully preserved, therefore a||XW1Y1|| * The term can be ignored, while for the l1 norm constraints of Y1 and Y2, l is used. 2,1Norm constraints are replaced to enhance feature selection capabilities. The final objective function of the joint sparse low-rank dual-dictionary learning model can be expressed as:

[0073]

[0074] Where a, b, c, and d are the balance parameters of the corresponding terms.

[0075] 2. Objective function optimization

[0076] For the objective function of the established model, this is a non-convex problem with respect to W1, W2, Y1, and Y2; however, it is a convex problem for any one of these variables. In this invention, an alternating optimization mechanism is used to solve the objective function (4). The idea of ​​the alternating iterative optimization mechanism is to fix the remaining variables and solve one variable. After obtaining the solution, the variable being solved is fixed, and the optimization problem of the previously fixed variable is solved until a relatively stable local optimum is obtained. In order to solve the objective function, Lagrange multipliers are introduced to constrain the variables W1, W2, Y1, and Y2, and constraint terms are added. Construct the Lagrange function:

[0077]

[0078] Where η and μ are penalty parameters, for ||A|| 2,1 Introducing a diagonal matrix U ii = 1 / 2||A||2, and initialize it as an identity matrix. Then ||A|| 2,1 Transform into matrix trace Tr(A) T The form is UA). For ||A||1, introduce two all-1 vectors i. c and i d Transform ||A||1 into Tr(i) c T A T i d For the nuclear norm ||XW1|| * Its physical meaning can be interpreted as the sum of the singular values ​​of the matrix, therefore we have ||XW1|| * =Tr(Σ), and define XW1 = MΣN T , and are the left singular matrix and the right singular matrix, respectively, and both are orthogonal matrices.

[0079] Referring to the above method and the transformation relationship between norm and matrix trace, equation (5) can be transformed into the form of matrix trace:

[0080]

[0081] U1 and U2 represent Y1 and Y2 respectively as l 2,1 The auxiliary diagonal matrix introduced by arithmetic operations is derived from the following relationship:

[0082]

[0083] Then we begin updating each variable. First, we fix W2, Y1, and Y2, update matrix W1, and calculate the partial derivatives of equation (6).

[0084]

[0085] Use KKT optimization conditions to let

[0086]

[0087] achievable

[0088]

[0089] Further, the update and iteration rules of W1 are obtained.

[0090]

[0091] Similarly, the update rules for W2, Y1, and Y2 can be obtained. The partial derivatives of the objective function with respect to variables W2, Y1, and Y2 are calculated as follows:

[0092]

[0093] Use KKT optimization conditions to let Further, the update rules for W2, Y1, and Y2 are obtained.

[0094]

[0095] 3. Convergence conditions

[0096] The maximum number of iterations is set to 500, and the stopping criterion is defined as follows:

[0097]

[0098] Where D1 = XW1, D2 = XW2, when RelErr is 10 -3 Stop updating the optimization process when necessary.

[0099] 4. Fault detection applications and criteria

[0100] After learning and obtaining the dictionary matrices D1 and D2, the test data X is... new And by reconstructing the coefficient matrices Y1 and Y2, we obtain and

[0101]

[0102] T 2Statistics measure the change of a sample vector in the principal space.

[0103]

[0104] To detect faults, if the test statistic exceeds the control limit, a fault is considered to have occurred; otherwise, no fault has occurred. Therefore, the test logic can be defined as follows:

[0105]

[0106] in, The control limit, representing a confidence level of α, is typically calculated using KDE (kernel density estimation).

[0107] To measure the performance of fault detection, a widely used metric is the alarm rate (FDR). The FDR is defined as follows:

[0108]

[0109] Where f = 0 means that no failure occurred in the test sample set, and f ≠ 0 means that a failure occurred in the test sample set.

[0110] Example 2:

[0111] This study uses the classic Tennessee Eastman (TE) chemical process simulation dataset as the research object. The TE process dataset is widely used in the field of industrial fault detection, containing complex dynamic chemical processes and covering core unit operations such as distillation and reaction, making it highly representative. The dataset provides 52 process variables (including 41 continuous variables and 11 discrete variables) and various common fault types (21 types in total, including 1 normal operating condition and 20 fault operating conditions). In this study, the normal data and 20 fault data provided by the TE process are selected as the analysis objects. Each fault starts at time 160 after the simulation data is generated and continues until time 960.

[0112] Samples under normal operating conditions are selected as the training set, and samples under fault conditions are selected as the test set X. new The experiment used the TE process dataset, selecting representative faults from the TE process for fault detection performance analysis, including fault 1 (IDV1), fault 2 (IDV2), fault 6 (IDV6), fault 7 (IDV7), fault 8 (IDV8), and fault 14 (IDV14). Comparisons were made with traditional PCA, ICA, and NMF fault detection methods. These faults have different characteristics, allowing for a comprehensive evaluation of the model's detection performance.

[0113] Table 1. Description of the 6 faults selected during the TE process.

[0114]

[0115]

[0116] The experimental results are shown in Table 2. Compared with the three traditional methods, the dual-dictionary learning method based on joint sparse low-rank representation achieves better performance in FDR. As shown in Table 2, in several representative fault types, the proposed method improves FDR by an average of 2.33% compared with PCA, 2.58% compared with ICA, and 2.77% compared with NMF. Therefore, the superiority of the proposed method is well-documented, and it also shows that combining sparse representation and low-rank representation with dual-dictionary learning is reasonable and effective.

[0117] Table 2 Fault Detection Rates (FDR) of Four Methods under Six Fault Types

[0118]

[0119] To make the fault detection results clearer, the detection process statistics for the selected fault types are plotted. Figure 2 As can be seen from the observation, the method proposed in this invention can detect the fault in all selected faults around the 160th sample, which confirms the satisfactory monitoring performance of this method.

[0120] The above specific embodiments are merely explanations of the present invention and are not intended to limit the present invention. After reading this specification, those skilled in the art can make modifications to these embodiments without contributing any inventive step, but as long as they are within the scope of the claims of the present invention, they are protected by patent law.

Claims

1. An industrial fault detection method based on joint sparse low-rank dual-dictionary learning, characterized by: The method includes the following steps; S1. Establish the objective function model: Sub-dictionaries D1 and D2 are established using a single-dictionary model. By collaboratively learning D1 and D2, a dual-dictionary learning model is formed. S2. Optimize the objective function model: Introduce Lagrange multipliers to constrain the variables in the objective function model, and use an alternating iterative optimization mechanism to optimize the objective function model; S3. Fault Detection Applications and Criteria: When the test statistic T... 2 If the control limit is exceeded, it is determined that a fault has occurred; otherwise, no fault has occurred.

2. The industrial fault detection method based on joint sparse low-rank dual-dictionary learning as described in claim 1, characterized in that: In S1, Y1 and Y2 are the encoding matrices corresponding to the two sub-dictionaries being learned, D1 and Y1 represent the low-rank dictionary learning part, and D2 and Y2 represent the sparse representation dictionary learning part. The nuclear norm is ||·|| * The structure is constrained to be low-rank, and the structure is constrained to be sparse by the l1 norm. a, b, c, and d are the balance parameters of the corresponding terms.

3. The industrial fault detection method based on joint sparse low-rank dual dictionary learning as described in claim 2, characterized in that: The sub-dictionary in S1 is represented as the product of the original data of the industrial process and the dictionary composition matrix, and l is used 2,1 Replacing the l1 norm constraints of Y1 and Y2 with norm constraints, the final objective function model is:

4. The industrial fault detection method based on joint sparse low-rank dual dictionary learning as described in claim 3, characterized in that: S2 introduces Lagrange multipliers to constrain variables W1, W2, Y1, and Y2, constructing the Lagrange function based on the final objective function model: η and μ are penalty parameters.

5. The industrial fault detection method based on joint sparse low-rank dual dictionary learning as described in claim 4, characterized in that: in Introducing a diagonal matrix U into the Lagrange function ii =1 / 2||A||2, change ||A|| 2,1 Transformed into the matrix trace Tr(A) T In the form of UA), two all-1 vectors i are introduced. c and i d Transform ||A||1 into Tr(i) c T A T i d ), thus obtaining the final Lagrange function:

6. The industrial fault detection method based on joint sparse low-rank dual dictionary learning as described in claim 5, characterized in that: Using an alternating iterative optimization mechanism with the final Lagrangian function as a template and employing KKT optimization conditions, W1, W2, Y1, and Y2 are optimized sequentially to obtain the W1 update iteration rule: Obtain the W2 update iteration rules: The Y1 update iteration rule is obtained as follows: The Y2 update iteration rule is obtained:

7. The industrial fault detection method based on joint sparse low-rank dual dictionary learning as described in claim 6, characterized in that: The KKT optimization conditions are as follows:

8. The industrial fault detection method based on joint sparse low-rank dual dictionary learning as described in claim 6, characterized in that: The convergence condition of the objective function model is: when RelErr is 10... -3 The optimization process is stopped when the time is right. The RelErr model is as follows: Where D1 = XW1, D2 = XW2.

9. The industrial fault detection method based on joint sparse low-rank dual dictionary learning as described in claim 8, characterized in that: The number of iterations is ≤500.

10. The industrial fault detection method based on joint sparse low-rank dual dictionary learning as described in claim 1, characterized in that: The specific steps of S3 are as follows: test data X new And by reconstructing the coefficient matrices Y1 and Y2, we obtain and Thus, T is obtained. 2 Statistics measure the change of a sample vector in the principal component space: The fault detection verification logic is as follows: The control limit represents the confidence level α, when the test statistic T 2 If the control limit is exceeded, it is determined that a fault has occurred; otherwise, no fault has occurred.