Control theory guided machine learning process modeling and fault detection method and system

CN121279085BActive Publication Date: 2026-08-07UNIV OF SCI & TECH BEIJING
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511347052.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2026-08-07
Estimated Expiration
2045-09-19

AI Technical Summary

Technical Problem

[0005]针对现有的基于数据驱动或神经网络的工业过程建模与故障检测技术模型准确率低或可解释性差以及故障检测方案实施较为困难等问题,本发明的目的在于提供一种控制理论引导机器学习(Control Theory Informed Machine Learning, CTIML)的过程建模与故障检测方法及系统

Benefits of technology

本发明提出了一种面向复杂工业过程中的一般非线性系统的过程建模与故障检测方案,融合了神经网络技术、稳定的核表征和控制理论知识,不但建立了精确的过程模型,有效学习和提取了过程运行动态,而且能够及时检测出过程运行中的异常,保障了过程运行的安全,该方案仅依赖工业过程运行中积累的数据,易于实现且工程成本低。并且,本发明创新性地将控制理论知识融入神经网络的训练和学习过程中,以指导和优化神经网络的建模效果,实验结果证明该策略在一定程度上能更充分地提取到过程运行数据中蕴含的动态,同时提升了模型的可解释性。此外,本发明适用于一般的非线性过程(或系统),而且易于嵌入已知的过程机理等先验知识,即模型中相应的神经网络模块可以直接用已知的非线性函数替代,易于配置,灵活性高。总之,本发明提出的控制理论引导机器学习的过程建模与故障检测方法建模精准、检测效果优良,具有较强的应用价值。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121279085B_ABST
    Figure CN121279085B_ABST
Patent Text Reader

Abstract

The application discloses a process modeling and fault detection method and system for guiding machine learning by control theory, and the method comprises the following steps: acquiring historical operation data of an industrial process, extracting operation dynamics in the historical operation data by using a neural network, and approximating an observer gain part of a kernel representation model to realize complete kernel representation model design; constructing an adjoint system of the kernel representation model by using Hamilton system theory, and guiding and optimizing a learning and training process of the kernel representation model and the adjoint system by using lossless constraints to realize design of a regularized kernel representation model; constructing detection statistics by using residuals of the kernel representation model and the regularized kernel representation model, setting a maximum value of the detection statistics when the industrial process normally operates as a fault detection threshold, and thus realizing fault detection on the industrial process to guarantee safety of the industrial process. In addition, by designing an online fine-tuning strategy, the generalization and applicability of the model are improved, and the method has strong practical value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of modeling and fault detection technology for complex industrial processes, and in particular to a process modeling and fault detection method and system guided by control theory and machine learning. Background Technology

[0002] With the deep integration of new-generation information and computer technologies into industrial processes, today's complex industrial processes are becoming increasingly large-scale and structurally intricate. They often require the coordinated operation of numerous sensors and actuators to complete production and processing tasks. However, even a minor malfunction can disrupt the normal operation of an industrial process and even cause losses. To maintain the safe and stable operation of industrial processes, fault detection technology has emerged. Over the past few decades, driven by both technological advancements and the demands of the industrial manufacturing market, fault detection technology has been extensively researched and applied. Fault detection technologies are often categorized into model-based methods, knowledge-based methods, and signal-based methods. In comparison, model-based fault detection methods have gained favor in industrial process applications due to their high accuracy and interpretability. However, this method heavily relies on analytical models of industrial processes, and mechanistic analysis and modeling of industrial processes is often difficult or requires extremely high engineering costs, thus limiting its application in industrial processes.

[0003] Benefiting from the large amounts of process data generated during industrial processes, data-driven modeling techniques have developed rapidly. Traditional data-driven modeling techniques mainly include least squares identification and subspace identification. These techniques, starting entirely from the data perspective, utilize causal analysis and correlation analysis to construct a model representation paradigm comparable to mechanistic analysis modeling, making them easy to implement and cost-effective. However, the models established by these techniques primarily represent linear industrial processes, thus affecting their ability to handle nonlinear industrial processes. Building on this, fuzzy dynamic modeling techniques connect linear models established at multiple operation points using numerous "If-Then" rules, thus forming a nonlinear description of industrial processes. However, the difficulty in collecting large amounts of data from multiple operation points in actual industrial processes reduces the engineering practicality of this technique. Compared to traditional data-driven modeling techniques, machine learning techniques are gaining increasing favor among researchers due to their powerful nonlinear fitting and generalization capabilities. Machine learning techniques utilize several neurons to extract features contained in process data, continuously adjusting and updating the connection weights between neurons based on the deviation between predicted and actual values ​​to minimize model errors. However, machine learning models are essentially "black boxes," and their outputs lack interpretability, which has become a bottleneck for the application of this technology in industrial processes.

[0004] Furthermore, both traditional data-driven modeling techniques and neural network techniques are based on offline training using collected process data. If the distribution of online and offline process data differs, it becomes difficult to adjust model parameters online to adapt the model to the dynamic online state of the industrial process. Therefore, traditional data-driven or neural network-based industrial process modeling and fault detection techniques are no longer adequate for the complex application requirements of today's industrial processes. Summary of the Invention

[0005] To address the problems of low accuracy, poor interpretability, and difficulty in implementing fault detection schemes in existing data-driven or neural network-based industrial process modeling and fault detection technologies, this invention aims to provide a process modeling and fault detection method and system based on Control Theory Informed Machine Learning (CTIML). This invention combines neural network technology, stable kernel representation (SKR), and specific control theory knowledge, integrating neural network technology into a stable kernel representation model. Guided by control theory knowledge, the learning and training process of the neural network is optimized, enabling the model to continuously approximate the real dynamics of the process. Based on this, this invention not only enables fault detection in industrial processes, thus ensuring their safety, but also proposes an online fine-tuning strategy. By introducing a residual neural network (ResNet) structure for learning, it improves the model's adaptability to fluctuations occurring during normal operation of the industrial process, enhancing the model's generalization and applicability.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: On the one hand, a method for process modeling and fault detection guided by control theory and machine learning is provided, the method comprising the following steps: S1. Design a kernel representation model based on a neural network, including: acquiring historical operating data of an industrial process, using a neural network to extract the operating dynamics from the historical operating data on the one hand, and approximating the observer gain part of the kernel representation model on the other hand to achieve a complete kernel representation model design. S2. Design a regularized kernel representation model, including: based on the kernel representation model, constructing the adjoint system of the kernel representation model using Hamiltonian system theory, and using lossless constraints to guide and optimize the learning and training process of the kernel representation model and its adjoint system, thereby realizing the design of the regularized kernel representation model. S3. Fault detection based on kernel representation model and regularized kernel representation model, including: constructing a detection statistic using the residuals of the kernel representation model and the regularized kernel representation model, and setting the maximum value of the detection statistic when the industrial process is running normally as the fault detection threshold, thereby realizing fault detection of the industrial process operation; S4. Design an online fine-tuning strategy, including: based on the kernel representation model, introduce a residual neural network structure for learning. The original kernel representation model is used to learn the operational dynamics of the collected offline data, and the newly introduced residual neural network structure is used to learn the new operational dynamics contained in the online data of the industrial process, so as to realize the online fine-tuning of the kernel representation model.

[0007] Optionally, S1 specifically includes: S11. For actual industrial processes, the operational dynamics are represented as follows:

[0008] in , , These represent the state vector, input vector, and output vector of the industrial process, respectively. It is the derivative of the state vector. Represents a continuously differentiable nonlinear function; A state-space description representing an actual industrial process. The kernel characterization model is expressed as:

[0009] in , They represent the state vectors respectively and output vector The estimate, express The derivative of The model represents the output vector. The residual, Indicates observer gain feedback; S12. Use neural networks to approximate the unknown parts of the kernel representation model, that is:

[0010] in , and Represents a neural network module; S13, Based on residuals Configure a loss function to train and optimize the parameters of the neural network, wherein the loss function is set as follows:

[0011] in N This indicates the number of data points collected.

[0012] Optionally, S2 specifically includes: S21. The kernel representation model of the actual industrial process is equivalently represented as:

[0013] in , , ; Using Hamiltonian system theory, construct Accompanying system :

[0014] in , , These represent the co-state vector, input vector, and output vector of the adjoint system, respectively. express The derivative of Indicates partial derivative, subscript T Indicates transpose; S22, Connection and Hamilton System Represented as:

[0015] S23. Integrating the lossless theorem into the training and learning process of neural networks, the lossless theorem is expressed as:

[0016]

[0017] With the help of neural network modules to approach The design of a regularized kernel representation model is achieved by adding additional constraints to all neural network modules using the lossless theorem. S24. The loss function of the regularized kernel representation model is set as follows:

[0018] in It is the first loss function used to ensure the model's tracking ability. It is used to optimize neural network modules. At the same time, it ensures that the kernel representation model is a regularized second loss function. It is the third loss function used to constrain the initial conditions of the energy function. It is used to satisfy Harmony state variables The fourth loss function relating the partial derivatives between them. , , , These are the weights of each loss function term.

[0019] Optionally, S3 specifically includes: S31. Constructing the detection statistic using the residuals of the kernel representation model and the regularized kernel representation model. J The residual sum of squares over a predetermined time period is used to construct the detection statistic to avoid the influence of excessively anomalous residuals at a single moment. Its calculation expression is as follows:

[0020] in Indicates the size of the sliding window; S32. When the system is in normal operating condition, the detection statistics... J It should be less than the fault detection threshold. J th Therefore, the maximum value of the detection statistic during normal operation of the industrial process is set as the fault detection threshold, i.e.:

[0021] in ; S33, By comparison J and J th The relationship between the magnitudes is used to determine whether a fault has occurred during the operation of the industrial process. The fault detection logic is set as follows:

[0022] This enables fault detection in industrial processes.

[0023] Optionally, S4 specifically includes: S41. Design an "online" kernel representation model based on residual neural networks, specifically expressed as follows:

[0024] in , ... represent the online fine-tuning residual neural network module; the "online" kernel representation model refers to a kernel representation model that includes the online fine-tuning residual neural network module; S42. Train the “online” kernel representation model offline using the pre-collected batch running data; S43. When the industrial process is running online, the parameters of the online fine-tuning residual neural network module are adjusted using the online gradient descent method, so that the "online" kernel representation model can continuously learn new knowledge contained in the running data. The update strategy is expressed as:

[0025] in and They represent t Time and t The parameters of the residual neural network module are fine-tuned online at time +1. Indicates the learning rate. This represents the gradient of the loss function.

[0026] On the other hand, a process modeling and fault detection system guided by control theory for machine learning is provided, for implementing the method described in any of the above embodiments, the system comprising: The kernel representation model design module is used to acquire historical operating data of industrial processes. It uses neural networks to extract the operating dynamics from the historical operating data and to approximate the observer gain part of the kernel representation model to achieve a complete kernel representation model design. The regularized kernel representation model design module is used to construct the adjoint system of the kernel representation model based on the kernel representation model using Hamiltonian system theory, and to guide and optimize the learning and training process of the kernel representation model and its adjoint system using lossless constraints, thereby realizing the design of the regularized kernel representation model. The fault detection module is used to construct a detection statistic using the residuals of the kernel representation model and the regularized kernel representation model, and set the maximum value of the detection statistic when the industrial process is running normally as the fault detection threshold, thereby realizing fault detection of the industrial process operation. The online fine-tuning module is used to introduce a residual neural network structure for learning based on the kernel representation model. The original kernel representation model is used to learn the dynamics of the collected offline data, while the newly introduced residual neural network structure is used to learn the new dynamics contained in the online data of the industrial process, thereby realizing the online fine-tuning of the kernel representation model.

[0027] On the other hand, an electronic device is provided, the electronic device comprising: processor; The memory stores computer-readable instructions, which, when loaded and executed by the processor, implement the steps of the process modeling and fault detection method for control theory-guided machine learning as described above.

[0028] On the other hand, a computer-readable storage medium is provided, wherein program code is stored in the computer-readable storage medium, and the program code can be called by a processor to execute the steps of the process modeling and fault detection method for machine learning guided by control theory as described above.

[0029] The beneficial effects of the technical solution provided by this invention include at least the following: This invention proposes a process modeling and fault detection scheme for general nonlinear systems in complex industrial processes. It integrates neural network technology, stable kernel representation, and control theory knowledge. This scheme not only establishes an accurate process model and effectively learns and extracts process dynamics, but also promptly detects anomalies during process operation, ensuring safe operation. Relying solely on data accumulated during industrial process operation, this scheme is easy to implement and has low engineering costs. Furthermore, this invention innovatively integrates control theory knowledge into the training and learning process of the neural network to guide and optimize its modeling performance. Experimental results demonstrate that this strategy can, to a certain extent, more fully extract the dynamics inherent in the process operation data, while improving the model's interpretability. In addition, this invention is applicable to general nonlinear processes (or systems) and easily incorporates known prior knowledge such as process mechanisms; that is, the corresponding neural network modules in the model can be directly replaced by known nonlinear functions, making it easy to configure and highly flexible. In summary, the control theory-guided machine learning process modeling and fault detection method proposed in this invention provides accurate modeling and excellent detection results, possessing strong application value. Attached Figure Description

[0030] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0031] Figure 1 This is a flowchart of a process modeling and fault detection method for machine learning guided by control theory, provided in an embodiment of the present invention. Figure 2 This is a technical roadmap for a process modeling and fault detection method guided by control theory in machine learning, provided by an embodiment of the present invention. Figure 3 This is a schematic diagram of the looper system provided in an embodiment of the present invention; Figure 4 This is a diagram of the actual operation data monitoring interface of the looper system in the factory provided in this embodiment of the invention; Figure 5 This is a graph showing the loss function during the training process of the kernel representation model and the regularized kernel representation model provided in this embodiment of the invention. Figure 6 This is a diagram showing the fitting effect of the subspace model, kernel representation model, and regularized kernel representation model provided in the embodiments of the present invention on the loop angle; Figure 7 This is a diagram showing the fitting effect of the subspace model, kernel characterization model, and regularized kernel characterization model provided in this embodiment of the invention on the strip tension. Figure 8 In the figures (a) and (b), the detection results of the kernel characterization model and the regularized kernel characterization model provided in the embodiments of the present invention on the tension sensor faults are respectively. Figure 9 In the figures (a) and (b), the detection results of the kernel representation model and the regularized kernel representation model provided in the embodiments of the present invention on the angle sensor fault are respectively. Figure 10 These are the prediction results of the kernel representation model and the "online" kernel representation model provided in this embodiment of the invention for the loop angle; Figure 11 This is a diagram showing the effect of the nuclear characterization model and the "online" nuclear characterization model provided in this embodiment of the invention on the prediction of strip tension; Figure 12 This is a schematic diagram of a process modeling and fault detection system for machine learning guided by control theory, provided in an embodiment of the present invention. Detailed Implementation

[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the described embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0033] This invention provides a method for process modeling and fault detection guided by control theory and machine learning, with reference to... Figure 1 and Figure 2 As shown, the method includes the following steps: S1. Design a kernel representation model based on a neural network, including: acquiring historical operating data of an industrial process, using a neural network to extract the operating dynamics from the historical operating data, and approximating the observer gain part of the kernel representation model to achieve a complete kernel representation model design.

[0034] As an optional embodiment of the present invention, S1 specifically includes: S11. For actual industrial processes, the operational dynamics are represented as follows:

[0035] in , , These represent the state vector, input vector, and output vector of the industrial process, respectively. It is the derivative of the state vector. Represents a continuously differentiable nonlinear function; A state-space description representing an actual industrial process. The kernel characterization model is expressed as:

[0036] in , They represent the state vectors respectively and output vector The estimate, express The derivative of The model represents the output vector. The residual, Indicates observer gain feedback; S12. Considering the lack of knowledge about the real dynamics of industrial processes and the difficulty in establishing observer gain feedback, a neural network is used to approximate the unknown parts of the kernel representation model, namely:

[0037] in , and Represents a neural network module; S13. In order to make the kernel characterization model as close as possible to the real dynamics of the industrial process, it can be based on the residuals. Configure a loss function to train and optimize the parameters of the neural network, wherein the loss function is set as follows:

[0038] in N This indicates the number of data points collected.

[0039] S2. Design a regularized kernel representation model, including: based on the kernel representation model, constructing the adjoint system of the kernel representation model using Hamiltonian system theory, and using lossless constraints to guide and optimize the learning and training process of the kernel representation model and its adjoint system, thereby realizing the design of the regularized kernel representation model.

[0040] As an optional embodiment of the present invention, S2 specifically includes: S21. The kernel representation model of the actual industrial process is equivalently represented as:

[0041] in , , ; Using Hamiltonian system theory, construct Accompanying system :

[0042] in , , These represent the co-state vector, input vector, and output vector of the adjoint system, respectively. express The derivative of Indicates partial derivative, subscript T Indicates transpose; S22, Connection and Hamilton System Represented as:

[0043] S23. Integrate the lossless property theorem into the training and learning process of neural networks. The lossless property theorem is expressed as:

[0044]

[0045] With the help of neural network modules to approach The design of a regularized kernel representation model is achieved by adding additional constraints to all neural network modules using the lossless theorem. S24. The loss function of the regularized kernel representation model is set as follows:

[0046] in It is the first loss function used to ensure the model's tracking ability. It is used to optimize neural network modules. At the same time, it ensures that the kernel representation model is a regularized second loss function. It is the third loss function used to constrain the initial conditions of the energy function. It is used to satisfy Harmony state variables The fourth loss function relating the partial derivatives between them. , , , These are the weights of each loss function term.

[0047] S3. Fault detection based on kernel representation model and regularized kernel representation model, including: constructing detection statistics using the residuals of the kernel representation model and the regularized kernel representation model, and setting the maximum value of the detection statistics when the industrial process is running normally as the fault detection threshold, thereby realizing fault detection of the industrial process operation.

[0048] As an optional embodiment of the present invention, S3 specifically includes: S31. Considering that the residuals of the kernel characterization model and the regularized kernel characterization model can reflect the operating status of the process to a certain extent, if an anomaly (failure) occurs during the process operation, the residuals will change significantly. Therefore, the residuals of the kernel characterization model and the regularized kernel characterization model can be used to construct a detection statistic. J The residual sum of squares over a predetermined time period is used to construct the detection statistic to avoid the influence of excessively anomalous residuals at a single moment. Its calculation expression is as follows:

[0049]

[0050] in Indicates the size of the sliding window; S32. When the system is in normal operating condition, the detection statistics... J It should be less than the fault detection threshold. J th Therefore, the maximum value of the detection statistic during normal operation of the industrial process is set as the fault detection threshold, i.e.:

[0051] in ; S33. Considering that fault detection is achieved through comparison... J and J th The relative sizes of these values ​​are used to determine whether a fault has occurred during the operation of the industrial process; therefore, the fault detection logic is set as follows:

[0052] This enables fault detection in industrial processes, thereby ensuring the safety and stability of process operation.

[0053] S4. Design an online fine-tuning strategy to improve the generalization and applicability of the model, including: on the basis of the kernel representation model, introduce a residual neural network structure for learning. The original kernel representation model is used to learn the dynamics of the collected offline data, and the newly introduced residual neural network structure is used to learn the new dynamics contained in the online data of the industrial process, so as to realize the online fine-tuning of the kernel representation model.

[0054] As an optional embodiment of the present invention, S4 specifically includes: S41. Considering that normal fluctuations such as production mode switching and setpoint changes may occur during industrial process operation, causing the established kernel representation model to no longer be accurately adapted, an "online" kernel representation model based on residual neural network (ResNet) is designed, specifically as follows:

[0055] in , ... represent the online fine-tuning residual neural network module; the "online" kernel representation model refers to a kernel representation model that includes the online fine-tuning residual neural network module; S42. The “online” kernel representation model is trained offline using the pre-collected batch running data, so that the model can extract as much process running dynamic knowledge as possible from the offline data. S43. When the industrial process is running online, the parameters of the online fine-tuning residual neural network module are adjusted using the Online Gradient Descent (OGD) method, enabling the "online" kernel representation model to continuously learn new knowledge contained in the running data. The update strategy is expressed as follows:

[0056] in and They represent t Time and t The parameters of the residual neural network module are fine-tuned online at time +1. The learning rate is typically represented by a small value or a value proportional to the number of hours the learning rate is given. , This represents the gradient of the loss function. It's understandable that here... t It can indicate the use of the first t The moment or the first t The first batch of data was processed. t Rounds of iteration are the biggest difference compared to offline training.

[0057] In this invention, control theory knowledge is integrated into the training and learning process of neural networks to optimize the modeling and monitoring of industrial processes, proposing a novel method for industrial process modeling and fault detection guided by control theory and machine learning. By combining neural network technology with stable kernel representations and optimizing the training and learning process of neural networks under the guidance of control theory, an accurate industrial process model is established. Furthermore, this model not only enables fault detection in industrial processes, ensuring their safe and stable operation, but also allows for online fine-tuning of model parameters, thereby improving the model's generalization and applicability. Therefore, this invention, by combining neural network technology, stable kernel representations, and specific control theory knowledge, has strong practical value for the accurate modeling and fault detection of industrial processes.

[0058] The effectiveness of the proposed control theory-guided machine learning-based process modeling and fault detection method will be verified below using the looper system in the hot strip rolling process. Figure 3 As shown, the looper system mainly regulates the looper angle and strip tension by adjusting the valve core displacement and the main frame speed to ensure smooth operation and avoid steel piling and pulling phenomena. Therefore, this invention considers the valve core displacement and main frame speed as input signals, and the looper angle and strip tension as output signals. Figure 4 As shown, this invention uses data from the looper system during actual production at a steel plant in 2023 for experiments. The rolling time for each slab is typically between 40 and 50 seconds. Considering that neural networks usually require a large amount of data for training and testing, this invention sets the sampling time to 0.01 seconds to acquire 10 sets of data from stable system operation, with approximately 3000 sampling points in each set. Furthermore, for fault detection purposes, this invention additionally collects two sets of data with similar operating conditions to the above data to add sensor faults and compensate for missing fault data. In addition, a set of data was collected showing fluctuations caused by changes in the setpoint during looper system operation to verify the "online" kernel representation model. The specific experimental results are as follows:

[0059] (1) Training and Implementation of Kernel Representation Model and Regularized Kernel Representation Model. To enable the kernel representation model and regularized kernel representation model to fully learn the operational dynamics contained in the operational data, this invention uses 10 sets of data collected during normal process operation for training and testing. To compare the training effects of the two, this invention records the changes in the loss function of the kernel representation model and the changes in the total loss function and residual loss term of the regularized kernel representation model during training, such as... Figure 5As shown, it is evident that both the loss function of the kernel representation model and the residual loss term of the regularized kernel representation model converge quickly, and the kernel representation model exhibits a smaller reconstruction error. To intuitively present the modeling effects of the two, a set of the aforementioned data was randomly selected for testing and compared synchronously with the model obtained from subspace identification, such as... Figure 6 and Figure 7 As shown, both the kernel representation model and the regularized kernel representation model clearly have excellent fitting effects on the live-loop process data, far exceeding the level of subspace identification modeling; on the other hand, the goodness-of-fit metric commonly used in process modeling was selected. R 2 The modeling accuracy of the kernel representation model was quantitatively compared with that of the mean squared error (MSE) evaluation index. As shown in Table 1, it is clear that the fitting accuracy of the kernel representation model is slightly higher, which is consistent with the situation reflected by the change of the loss function during training.

[0060] Table 1

[0061] (2) Fault Detection. Based on the fault detection logic and offline training results of the (regularized) kernel representation model adopted in this invention, the fault detection thresholds of the kernel representation model and the regularized kernel representation model are 1.2748 and 1.9533, respectively. Considering the lack of fault data, additional signals were added to the output signals of two additional sets of normal data to simulate sensor faults, namely: ① 0.225 MPa (approximately 3.6%) slowly varying tension sensor fault: a linearly increasing fault was introduced from the 600th data point until the fault reached 0.225 MPa at the 800th data point and remained constant. The fault detection results are as follows: Figure 8 As shown, Figure 8 Figures (a) and (b) show the detection results of the kernel characterization model and the regularized kernel characterization model for tension sensor faults, respectively; ② 0.6° (approximately 2.4%) slowly varying angle sensor fault: A linearly increasing fault is introduced from the 800th data point until the fault reaches 0.6° at the 1000th data point and remains constant. The fault detection results are as follows: Figure 9 As shown, Figure 9 Figures (a) and (b) show the detection results of the kernel representation model and the regularized kernel representation model for angle sensor faults, respectively. Clearly, combining the results of the two sets of fault detection, although the kernel representation model fits the data more accurately, its fault detection performance is not as good as that of the regularized kernel representation model. This not only reflects the "optimality" of the regularized kernel representation model in the least squares sense, but also, to some extent, illustrates that the regularized kernel representation modeling strategy, which incorporates control theory, improves the interpretability of the model.

[0062] (3) Implementation of the “online” kernel representation model. For simplicity, this invention only uses a kernel representation model with an added “ResNet” structure to demonstrate the effect, that is, the “online” kernel representation model only... This was used for online fine-tuning of the model. To demonstrate its online fine-tuning capability, a set of data showing fluctuations caused by changes in setpoints during process operation was used for testing; specifically, fluctuations in the rolling data occurred after the 950th sampling point. The experimental results are as follows: Figure 10 and Figure 11 As shown, this invention demonstrates that the online fine-tuning structure proposed in this invention can dynamically learn new knowledge contained in process data and has strong adaptability.

[0063] Accordingly, embodiments of the present invention also provide an industrial process modeling and fault detection system, such as... Figure 12 As shown, the system includes: The kernel representation model design module 201 is used to acquire historical operating data of industrial processes, and to use a neural network to extract the operating dynamics from the historical operating data on the one hand, and to approximate the observer gain part of the kernel representation model on the other hand to achieve the complete kernel representation model design. The regularized kernel representation model design module 202 is used to construct the adjoint system of the kernel representation model based on the kernel representation model using Hamiltonian system theory, and to guide and optimize the learning and training process of the kernel representation model and its adjoint system using lossless constraints, thereby realizing the design of the regularized kernel representation model. The fault detection module 203 is used to construct a detection statistic using the residuals of the kernel representation model and the regularized kernel representation model, and set the maximum value of the detection statistic when the industrial process is running normally as the fault detection threshold, thereby realizing fault detection of the industrial process operation. The online fine-tuning module 204 is used to introduce a residual neural network structure for learning based on the kernel representation model. The original kernel representation model is used to learn the dynamics of the collected offline data, and the newly introduced residual neural network structure is used to learn the new dynamics contained in the online data of the industrial process, so as to realize the online fine-tuning of the kernel representation model.

[0064] For ease of explanation, Figure 12 Only the main components of the system are shown. The system of this embodiment can be used to perform... Figure 1 The technical solutions of the method embodiments shown are similar in principle and in effect, and will not be described again here.

[0065] In an exemplary embodiment, the present invention also provides an electronic device, the electronic device comprising: processor; The memory stores computer-readable instructions, which, when loaded and executed by the processor, implement the steps of the process modeling and fault detection method for control theory-guided machine learning as described above.

[0066] In an exemplary embodiment, the present invention also provides a computer-readable storage medium storing at least one instruction, which is loaded and executed by a processor to implement the steps of the process modeling and fault detection method for control theory-guided machine learning described above. For example, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical data storage device, etc.

[0067] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0068] The use of terms such as "an embodiment," "an embodiment," "an exemplary embodiment," and "some embodiments" in the specification indicates that the described embodiment may include a specific feature, structure, or characteristic, but not every embodiment necessarily includes that specific feature, structure, or characteristic. Furthermore, when a specific feature, structure, or characteristic is described in connection with an embodiment, implementing such a feature, structure, or characteristic in conjunction with other embodiments (whether explicitly described or not) should be within the knowledge of those skilled in the art.

[0069] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0070] In the embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0071] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0072] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0073] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0074] This invention encompasses any substitutions, modifications, equivalent methods, and solutions made within the spirit and scope of this invention. To provide the public with a thorough understanding of this invention, specific details are described in detail in the following preferred embodiments; however, those skilled in the art will fully understand the invention even without these details. Furthermore, to avoid unnecessary misunderstanding of the essence of this invention, well-known methods, processes, procedures, components, and circuits are not described in detail.

[0075] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A process modeling and fault detection method guided by control theory and using machine learning, characterized in that, Includes the following steps: S1. Design a kernel representation model based on a neural network, including: acquiring historical operating data of an industrial process, using a neural network to extract the operating dynamics from the historical operating data on the one hand, and approximating the observer gain part of the kernel representation model on the other hand to achieve a complete kernel representation model design. S2. Design a regularized kernel representation model, including: based on the kernel representation model, constructing the adjoint system of the kernel representation model using Hamiltonian system theory, and using lossless constraints to guide and optimize the learning and training process of the kernel representation model and its adjoint system, thereby realizing the design of the regularized kernel representation model. S3. Fault detection based on kernel representation model and regularized kernel representation model, including: constructing a detection statistic using the residuals of the kernel representation model and the regularized kernel representation model, and setting the maximum value of the detection statistic when the industrial process is running normally as the fault detection threshold, thereby realizing fault detection of the industrial process operation; S4. Design an online fine-tuning strategy, including: based on the kernel representation model, introduce a residual neural network structure for learning. The original kernel representation model is used to learn the operational dynamics of the collected offline data, and the newly introduced residual neural network structure is used to learn the new operational dynamics contained in the online data of the industrial process, so as to realize the online fine-tuning of the kernel representation model.

2. The process modeling and fault detection method guided by control theory for machine learning according to claim 1, characterized in that, S1 specifically includes: S11. For actual industrial processes, the operational dynamics are represented as follows: ; in , , These represent the state vector, input vector, and output vector of the industrial process, respectively. It is the derivative of the state vector. Represents a continuously differentiable nonlinear function; A state-space description representing an actual industrial process. The kernel characterization model is expressed as: ; in , They represent the state vectors respectively and output vector The estimate, express The derivative, The model represents the output vector. The residual, Indicates observer gain feedback; S12. Use neural networks to approximate the unknown parts of the kernel representation model, that is: ; in , and Represents a neural network module; S13, Based on residuals Configure a loss function to train and optimize the parameters of the neural network, wherein the loss function is set as follows: ; in N This indicates the number of data points collected.

3. The process modeling and fault detection method guided by control theory for machine learning according to claim 2, characterized in that, S2 specifically includes: S21. The kernel representation model of the actual industrial process is equivalently represented as: ; in , , ; Using Hamiltonian system theory, construct Accompanying system : ; in , , These represent the co-state vector, input vector, and output vector of the adjoint system, respectively. express The derivative, Indicates partial derivative, subscript T Indicates transpose; S22, Connection and Hamilton System Represented as: ; S23. Integrating the lossless theorem into the training and learning process of neural networks, the lossless theorem is expressed as: ; ; With the help of neural network modules to approach The design of a regularized kernel representation model is achieved by adding additional constraints to all neural network modules using the lossless theorem. S24. The loss function of the regularized kernel representation model is set as follows: ; in It is the first loss function used to ensure the model's tracking ability. It is used to optimize neural network modules. At the same time, it ensures that the kernel representation model is a regularized second loss function. It is the third loss function used to constrain the initial conditions of the energy function. It is used to satisfy Harmony state variables The fourth loss function relating the partial derivatives between them. , , , These are the weights of each loss function term.

4. The process modeling and fault detection method guided by control theory for machine learning according to claim 1, characterized in that, S3 specifically includes: S31. Constructing the detection statistic using the residuals of the kernel representation model and the regularized kernel representation model. J The residual sum of squares over a predetermined time period is used to construct the detection statistic to avoid the influence of residual anomalies at a single moment. Its calculation expression is as follows: ; in Indicates the size of the sliding window; S32. When the system is in normal operating condition, the detection statistics... J It should be less than the fault detection threshold. J th Therefore, the maximum value of the detection statistic during normal operation of the industrial process is set as the fault detection threshold, i.e.: ; in ; S33, By comparison J and J th The relationship between the magnitudes is used to determine whether a fault has occurred during the operation of the industrial process. The fault detection logic is set as follows: ; This enables fault detection in industrial processes.

5. The process modeling and fault detection method guided by control theory for machine learning according to claim 3, characterized in that, S4 specifically includes: S41. Design an "online" kernel representation model based on residual neural networks, specifically expressed as follows: ; in , ... represents the online fine-tuning residual neural network module; the "online" kernel representation model refers to a kernel representation model that includes the online fine-tuning residual neural network module; S42. Train the "online" kernel representation model offline using the pre-collected batch running data; S43. When the industrial process is running online, the parameters of the online fine-tuning residual neural network module are adjusted using the online gradient descent method, so that the "online" kernel representation model can continuously learn new knowledge contained in the running data. The update strategy is expressed as: ; in and They represent t Time and t The parameters of the residual neural network module are fine-tuned online at time +1. Indicates the learning rate. This represents the gradient of the loss function.

6. A process modeling and fault detection system guided by control theory for machine learning, the system being used to implement the method as described in any one of claims 1 to 5, characterized in that, The system includes: The kernel representation model design module is used to acquire historical operating data of industrial processes. It uses neural networks to extract the operating dynamics from the historical operating data and to approximate the observer gain part of the kernel representation model to achieve a complete kernel representation model design. The regularized kernel representation model design module is used to construct the adjoint system of the kernel representation model based on the kernel representation model using Hamiltonian system theory, and to guide and optimize the learning and training process of the kernel representation model and its adjoint system using lossless constraints, thereby realizing the design of the regularized kernel representation model. The fault detection module is used to construct a detection statistic using the residuals of the kernel representation model and the regularized kernel representation model, and set the maximum value of the detection statistic when the industrial process is running normally as the fault detection threshold, thereby realizing fault detection of the industrial process operation. The online fine-tuning module is used to introduce a residual neural network structure for learning based on the kernel representation model. The original kernel representation model is used to learn the dynamics of the collected offline data, while the newly introduced residual neural network structure is used to learn the new dynamics contained in the online data of the industrial process, thereby realizing the online fine-tuning of the kernel representation model.

7. An electronic device, characterized in that, The electronic device includes: processor; A memory storing computer-readable instructions that, when loaded and executed by the processor, implement the method as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program code that can be invoked by a processor to execute the method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Industrial process fault detection method based on prior information auxiliary neural network modeling

    CN118502389A

  • System modeling and fault detection method and system for strip steel hot continuous rolling loop

    CN119702715A