Chemical process fault detection method and device based on progressive decoupling representation learning network

By constructing a progressively decoupled representation learning network of real-time and historical data matrices, eliminating temporal correlations and calculating latent variables, the applicability problem of fault detection in complex nonlinear industrial processes in existing technologies is solved, and early high-sensitivity detection of minor faults is achieved.

CN121859178APending Publication Date: 2026-04-14BEIJING UNIV OF CHEM TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing multivariate statistical fault detection techniques are not sufficiently applicable to complex nonlinear industrial processes and cannot effectively capture multi-scale and multi-level complex patterns, resulting in insufficient early detection capabilities for weak and gradual faults.

Method used

A method based on progressive decoupling representation learning network is adopted. By constructing a real-time data matrix and a historical data matrix, the progressive temporal decoupling module is used to eliminate temporal correlation, and the progressive feature representation module is used to calculate latent variables and reconstruct outputs to form online fault monitoring statistics.

Benefits of technology

It maintains a stable statistical discrimination basis under dynamic time-varying and nonlinear operating conditions, improving the early detection sensitivity and alarm reliability of weak faults and gradual faults.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121859178A_ABST
    Figure CN121859178A_ABST
Patent Text Reader

Abstract

The invention discloses a chemical process fault detection method and device based on a progressive decoupling representation learning network, and the method comprises the steps: introducing a pre-trained progressive time sequence decoupling module for a dynamic time-varying and high-nonlinearity chemical process, eliminating time sequence correlation of process variables layer by layer through multi-level recursive orthogonal projection and nonlinear mapping to obtain standardized depth time uncorrelated components; and further performing deep and multi-level latent variable extraction and reconstruction modeling by using a pre-trained progressive feature representation module, and constructing online fault monitoring statistics based on each layer of latent variable and reconstruction output to realize online out-of-limit judgment and alarm under a stable control limit. Therefore, the characterization capability of a multi-scale and multi-level complex mode is improved, and the early detection sensitivity and the alarm reliability of weak faults and gradual faults can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of chemical process monitoring technology, and in particular to a method and equipment for chemical process fault detection based on a progressive decoupled characterization learning network. Background Technology

[0002] With the rapid development of modern industrial systems, industrial process fault detection has become crucial in ensuring operational safety, product quality, and economic efficiency. Due to the inherent high dimensionality and non-Gaussianity of contemporary process variables, multivariate statistical fault detection techniques have been widely applied. However, most multivariate statistical fault detection techniques are based on the fundamental assumptions of linear mixed models, which severely limits their applicability and effectiveness in complex nonlinear industrial processes.

[0003] To overcome this technological bottleneck, recent industrial research has gradually shifted towards combining the powerful nonlinear modeling capabilities of neural networks with the statistical independence principle of multivariate statistical fault detection. However, existing methods still reveal many fundamental technical challenges when facing practical industrial applications. First, existing neural network methods based on multivariate statistical fault detection rely on the assumption of time stationarity and static historical data for modeling, which fundamentally contradicts the dynamic and time-varying nature of actual industrial processes. Second, current methods have significant shortcomings in feature extraction depth, remaining at a shallow network architecture and failing to effectively capture the multi-scale, multi-level complex patterns prevalent in industrial processes, severely impacting the system's early detection capabilities for weak and gradual faults. Summary of the Invention

[0004] This application provides a method, system, device, storage medium, and program product for detecting faults in chemical processes based on a progressively decoupled characterization learning network, which is used to solve at least one of the above-mentioned technical problems.

[0005] In a first aspect, embodiments of this application provide a chemical process fault detection method based on a progressively decoupled representation learning network, comprising: acquiring real-time sampling data of the chemical process, and constructing a real-time data matrix containing current time information and a historical data matrix containing information of a preset number of past times based on the real-time sampling data; inputting the real-time data matrix and the historical data matrix into a pre-trained progressive temporal decoupling module, and using the progressive temporal decoupling module to perform multi-level recursive orthogonal projection and nonlinear mapping on the input data to eliminate temporal correlation and output standardized deep temporally uncorrelated components; inputting the deep temporally uncorrelated components into a pre-trained progressive feature representation module, and using the progressive feature representation module to calculate the latent variables and reconstructed outputs of each level of the network, and calculating the corresponding online fault monitoring statistics based on the latent variables and reconstructed outputs; and determining that the chemical process is in a fault state and triggering an alarm when the online fault monitoring statistics exceed a preset control limit.

[0006] Secondly, embodiments of this application provide a storage medium storing one or more programs including execution instructions, which can be read and executed by electronic devices (including but not limited to computers, servers, or network devices) to execute any of the above-mentioned chemical process fault detection methods based on progressive decoupling characterization learning networks.

[0007] Thirdly, a computer device is provided, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to execute any of the above-described chemical process fault detection methods based on progressively decoupled characterization learning networks of this application.

[0008] Fourthly, embodiments of this application also provide a computer program product, the computer program product including a computer program stored on a storage medium, the computer program including program instructions, which, when executed by a computer, cause the computer to execute any of the above-mentioned chemical process fault detection methods based on progressive decoupling representation learning networks.

[0009] The beneficial effects of the embodiments of this application are as follows: A chemical process fault detection method based on a progressively decoupled representation learning network simultaneously constructs a real-time data matrix containing current information and a historical data matrix containing information from a predetermined number of past moments. These two matrices are then input into a pre-trained progressive temporal decoupling module. Multi-level recursive orthogonal projection and nonlinear mapping are used to eliminate temporal correlations in the process data, resulting in standardized deep temporally uncorrelated components. These deep temporally uncorrelated components are further input into a pre-trained progressive feature representation module to calculate latent variables at each level and reconstruct the output, thereby forming online fault monitoring statistics. This allows fault determination to maintain a stable statistical basis under dynamic, time-varying, and nonlinear operating conditions, and improves the representation ability of multi-scale and multi-level complex patterns. This, in turn, helps improve the early detection sensitivity and alarm reliability of weak and gradual faults. Attached Figure Description

[0010] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 A flowchart illustrating an example of a chemical process fault detection method based on a progressively decoupled representation learning network according to an embodiment of this application is shown. Figure 2 A flowchart illustrating another example of a chemical process fault detection method based on a progressively decoupled representation learning network according to an embodiment of this application is shown. Figure 3 A schematic diagram of the deep neural network used by the progressive temporal decoupling module and the progressive feature representation module according to embodiments of this application is shown. Figure 4 The diagram shows the detection effect of the IDV (1) fault in the Tennessee Eastman process according to an embodiment of this application; Figure 5 The diagram shows the detection effect of the IDV (2) fault in the Tennessee Eastman process according to an embodiment of this application; Figure 6 The diagram shows the detection effect of the IDV (3) fault in the Tennessee Eastman process according to an embodiment of this application; Figure 7 The diagram shows the detection effect of the IDV (4) fault in the Tennessee Eastman process according to an embodiment of this application; Figure 8The diagram shows the detection effect of the IDV (5) fault in the Tennessee Eastman process according to an embodiment of this application; Figure 9 The diagram shows the detection effect of the IDV (6) fault in the Tennessee Eastman process according to an embodiment of this application; Figure 10 The diagram shows the detection effect of the IDV (7) fault in the Tennessee Eastman process according to an embodiment of this application; Figure 11 The diagram shows the detection effect of the IDV (8) fault in the Tennessee Eastman process according to an embodiment of this application; Figure 12 The diagram shows the detection effect of the IDV (9) fault in the Tennessee Eastman process according to an embodiment of this application; Figure 13 The diagram shows the detection effect of the IDV (10) fault in the Tennessee Eastman process according to an embodiment of this application; Figure 14 The diagram shows the detection effect of the IDV (11) fault in the Tennessee Eastman process according to an embodiment of this application; Figure 15 The diagram shows the detection effect of the IDV (12) fault in the Tennessee Eastman process according to an embodiment of this application; Figure 16 The diagram shows the detection effect of the IDV (13) fault in the Tennessee Eastman process according to an embodiment of this application; Figure 17 The diagram shows the detection effect of the IDV (14) fault in the Tennessee Eastman process according to an embodiment of this application; Figure 18 The diagram shows the detection effect of the IDV (15) fault in the Tennessee Eastman process according to an embodiment of this application; Figure 19 The diagram shows the detection effect of the IDV (16) fault in the Tennessee Eastman process according to an embodiment of this application; Figure 20 The diagram shows the detection effect of the IDV (17) fault in the Tennessee Eastman process according to an embodiment of this application; Figure 21The diagram shows the detection effect of the IDV (18) fault in the Tennessee Eastman process according to an embodiment of this application; Figure 22 The diagram shows the detection effect of the IDV (19) fault in the Tennessee Eastman process according to an embodiment of this application; Figure 23 The diagram shows the detection effect of the IDV (20) fault in the Tennessee Eastman process according to an embodiment of this application; Figure 24 This is a schematic diagram of the structure of an embodiment of the electronic device of this application. Detailed Implementation

[0012] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other.

[0013] It should also be noted that, in this document, the terms "comprising" or "including" include not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0014] Figure 1 A flowchart illustrating an example of a chemical process fault detection method based on a progressively decoupled representation learning network according to an embodiment of this application is shown.

[0015] Regarding the execution subject of the method in the embodiments of this application, it can be any controller or processor with computing or processing capabilities. In some examples, the method in the embodiments of this application can be integrated and configured in an electronic device or terminal through software, hardware or a combination of software and hardware, and the type of terminal or electronic device can be diverse, such as mobile phone, tablet computer, desktop computer or vehicle terminal, etc.

[0016] For example, the implementing entity of the method in this application embodiment can be a chemical process monitoring platform, employing a chemical process micro-fault detection method based on a progressive decoupling representation learning network. This method achieves highly sensitive detection of chemical process faults through cascaded processing of two stages: progressive temporal decoupling and progressive feature representation. Thus, a deep feature learning method capable of simultaneously handling temporal dynamic correlations and nonlinear features can improve the accuracy and sensitivity of chemical process fault detection.

[0017] Figure 1 A flowchart illustrating an example of a chemical process fault detection method based on a progressively decoupled representation learning network according to an embodiment of this application is shown.

[0018] like Figure 1 As shown, in step S110, real-time sampling data of the chemical process is acquired, and a real-time data matrix containing current time information and a historical data matrix containing a preset number of past time information are constructed based on the real-time sampling data.

[0019] In some implementations, data is collected from DCS / PLC / SCADA or industrial data platforms according to the sampling period. Δt (For example, 1 s, 5 s, or 1 min, determined based on device inertia and control cycle) Acquire multivariate real-time sampling data. Sampling variables may include key manipulated variables, controlled variables, and disturbance variables, such as reactor jacket temperature, kettle temperature, top / bottom temperature, pressure at each stage, feed / reflux / steam flow rate, liquid level, online analysis composition, etc. Preferably, to ensure the stability and comparability of subsequent network inputs, basic data processing can be performed after sampling. For example, forward hold or linear interpolation can be used for missing points, median filtering / amplification strategies can be used for obvious outliers in sensor spikes, variables of different dimensions can be standardized (e.g., z-score normalization using the mean and standard deviation obtained from the normal operating condition training set), and unusable channels can be shielded or weighted attenuation according to process constraints, thereby forming a consistent input that can be used for online calculation.

[0020] Based on the aforementioned real-time sampling vectors, two types of matrices are constructed to explicitly represent dynamic information: one is a real-time data matrix containing information at the current moment, which can be a single-moment multivariate row vector (1× m ) or short batch window ( b × m The second is a historical data matrix containing information about a predetermined number of past moments, typically using a length of [missing information]. L The sliding window will t -1 to t - L Multivariate data stacked in chronological order L × mThe matrix (which may further include multiple lags or difference terms to characterize the rate of change) is strictly aligned in time with the historical matrix (corresponding to the same "current moment"). t (within the context of ""), and updated on a rolling basis with a fixed refresh strategy, so that each detection carries information of "current state + recent evolution trajectory". In this way, process dynamics and operating condition drift can be directly perceived in the online stage, reducing the probability of misjudging normal dynamic changes as faults, and improving the adaptability and monitoring stability to time-varying operating conditions.

[0021] In step S120, the real-time data matrix and the historical data matrix are input into the pre-trained progressive temporal decoupling module. The progressive temporal decoupling module performs multi-level recursive orthogonal projection and nonlinear mapping on the input data to eliminate temporal correlation and output standardized deep temporally uncorrelated components.

[0022] In some implementations, the historical matrix is ​​expanded along the time dimension and concatenated with the real-time vector, or a dual-branch structure is used to encode the "current branch" and the "historical branch" separately before merging. The progressive temporal decoupling module can be composed of multiple levels of decoupling units connected in series. Each level of decoupling unit includes a recursive orthogonal projection operator and a nonlinear mapping network. The recursive orthogonal projection operator is used to remove components related to historical lag information layer by layer in the feature space. For example, in the l-th level, a set of basis vectors representing the lag subspace is first extracted from the historical branch (which can be achieved through QR decomposition / Gram-Schmidt orthogonalization or a learnable orthogonal constraint projection matrix). Then, the current feature is projected onto the orthogonal complement of the subspace, thereby obtaining residual features that are as orthogonal as possible to the lag subspace. In addition, the above process is performed recursively layer by layer, so that the "temporally related parts that can be linearly decoupled" are compressed and stripped away layer by layer. Subsequently, nonlinear mapping networks (such as multilayer perceptrons with residual connections, gating units, or lightweight convolutional / attention structures) perform nonlinear transformations on the projection residuals to handle the nonlinear dynamic coupling that is common in chemical processes, thereby further weakening the time-series dependencies that cannot be completely eliminated by linear orthogonal projection.

[0023] Furthermore, to achieve the output of deep time-independent components, the module's pre-training phase can introduce a time-decorrelation objective on normal operating data. For example, minimizing the autocorrelation / cross-correlation between the output features and their several lagged versions, or constraining the output covariance matrix to be close to a diagonal matrix, and using a standardization layer (batch normalization / layer normalization and whitening based on training set statistics) to output standardized components. Thus, when running online, the decoupling module can directly output standardized deep time-independent components.

[0024] In step S130, the deep time-independent components are input into the pre-trained progressive feature representation module. The progressive feature representation module is used to calculate the latent variables and reconstructed outputs of each layer of the network, and the corresponding online fault monitoring statistics are calculated based on the latent variables and reconstructed outputs.

[0025] In some implementations, the progressive feature representation module can be implemented as a progressive deep autoencoder structure, such as a multi-level encoder that extracts latent variables layer by layer from shallow to deep (lower layers characterize local and rapidly changing patterns, while higher layers characterize global and coupled structural patterns), and the corresponding decoder generates the reconstructed output.

[0026] In practice, each level outputs a set of latent variables. and corresponding reconstruction At the same time, the residual is obtained. To avoid the problem that shallow networks struggle to capture complex multi-scale patterns, progressive structures can employ a strategy of "increasing representational power layer by layer" (e.g., increasing hidden layer width / nonlinear depth layer by layer or introducing cross-layer jump connections). Through hierarchical consistency constraints, latent variables in each layer can maintain a compact representation of normal patterns while also generating observable deviations when anomalies occur.

[0027] Online fault monitoring statistics are calculated based on latent variables at each level and reconstructed residuals. A combination of methods commonly used in multivariate statistical monitoring and easily implemented in engineering can be adopted. Statistics are constructed for latent variables (using covariance estimates of latent variables under normal operating conditions or their diagonal approximations to form weighted distances), and SPE / Q statistics (residual energy or weighted residual energy) are constructed for reconstructed residuals. The total monitoring index is obtained by weighted fusion of statistics from different levels. The reference distribution required for control limits can be obtained from normal data in the offline stage using empirical quantile methods or nonparametric estimation to adapt to non-Gaussian characteristics; simultaneously, moving mean / exponential smoothing can be applied to online statistics to suppress single-point noise.

[0028] In step S140, if the online fault monitoring statistics exceed the preset control limit, the chemical process is determined to be in a fault state and an alarm is triggered.

[0029] In some implementations, to improve engineering practicality, "continuous limit exceedance" or "multiple limit exceedance" strategies can be used to suppress instantaneous false alarms caused by random noise. For example, an alarm can only be triggered if the statistical quantity exceeds the control limit for K consecutive sampling points (K can be set according to the device dynamics and alarm sensitivity requirements), or a dual threshold strategy with hysteresis can be used (triggered when exceeding the upper limit, and deactivated when returning to the lower limit) to reduce alarm jitter. At the same time, information such as key variables, latent variables, and residual contributions at the time of alarm can be packaged and recorded to facilitate operators in quickly locating the source of the anomaly and taking appropriate measures.

[0030] Because the statistics are based on time-independent components and multi-level latent variables / reconstructed residuals, their distribution stability and anomaly sensitivity are higher. Therefore, the control limit setting is more reliable, and the alarm triggering is more in line with the actual operating condition change pattern. With the help of online strategies such as continuous limit exceedance, the frequency of false alarms can be reduced without sacrificing early detection capabilities, thereby improving the robustness and maintainability of online monitoring of chemical plants and making the alarm more reliable and operable.

[0031] Regarding the implementation details of step S110, in some examples of embodiments of this application, based on a preset historical time series length, for each real-time sampling moment, process variable data from a preset number of consecutive sampling moments preceding that sampling moment are obtained. The process variable data from the preset number of consecutive sampling moments are concatenated in chronological order to form a historical information vector that can characterize the dynamic characteristics of the process. A real-time data matrix is ​​constructed based on the historical information vector of the current moment, and a historical data matrix is ​​constructed based on the historical information vectors of the preset number of past moments.

[0032] Specifically, considering that chemical process data often exhibits strong time autocorrelation, a single observation at a single moment is insufficient to fully reflect the current state and evolution trend of the system. Therefore, when collecting data containing... One sample After processing variable data for each variable, the system not only utilizes the current moment... observation vector Instead, it introduces a preset historical time series length. (e.g., lag order). For each sampling time... The system will backtrack and retrieve continuous data up to that point in time. The sampled data at each time point (i.e. The current data is then concatenated with these historical data in chronological order to construct an expanded-dimensional historical information vector. The vector It can be Dimensional Its dimensions are expanded to By "folding" information from the time dimension into the spatial dimension, it is possible to simultaneously model and handle the coupling relationships between variables and the dynamic dependencies in time within an augmented feature space, thereby effectively solving the problem that traditional static models cannot describe dynamic chemical processes.

[0033] Building upon this, the system further performs structured construction of the real-time data matrix and the historical data matrix to accommodate subsequent progressive time-series decoupling operations. Specifically, the system uses the historical information vector of each sample constructed above. The historical data is stacked as row or column vectors to form a historical data matrix for training or online analysis. For example, considering that the dynamic response characteristics of most chemical processes (such as reactor temperature, liquid level changes, etc.) can be approximated as a second-order inertial system in terms of physical mechanism, the historical time series length is... The value is set to 2. This is sufficient to effectively capture the main dynamic characteristics of the process (i.e., the current state is significantly related to the previous two states), while avoiding excessively high data matrix dimensions due to the introduction of excessively long historical time series. Thus, while ensuring the sensitivity of fault detection, the computational burden and model complexity are effectively controlled.

[0034] In some examples of embodiments of this application, before performing multi-level recursive orthogonal projection and nonlinear mapping on the input data using the progressive temporal decoupling module, the row full rank determination is performed on the real-time data matrix and the historical data matrix.

[0035] It should be noted that before inputting the constructed real-time data matrix and historical data matrix into the asymptotic time-series decoupling module for orthogonal projection, it is essential to ensure that the data meets specific algebraic conditions. Because real-world chemical process variables often exhibit extremely strong linear correlations (i.e., multicollinearity), directly constructed data matrices are highly prone to having rows with incomplete rank. This will result in the absence or highly unstable value of the inverse matrix required for subsequent orthogonal projection calculations.

[0036] Therefore, the historical data matrix is ​​first subjected to a row rank determination. If the matrix is ​​of full rank, subsequent decoupling is performed directly; if it is of less than rank, an adaptive regularization mechanism is activated. This mechanism aims to disrupt the complete collinearity of the data by injecting a small amount of statistically controllable random perturbation into the original data, thereby forcibly restoring the invertibility of the matrix without changing the main characteristics of the data.

[0037] When the determination matrix is ​​not of full rank, the relative coefficients of the control noise intensity are calculated, and independent Gaussian noise is introduced into the real-time data matrix and the historical data matrix based on the relative coefficients to perform adaptive regularization processing, thereby ensuring the orthogonal projection invertibility of matrix operations.

[0038] Specifically, adaptive regularization does not simply add white noise, but rather injects customized noise based on the statistical distribution characteristics of each variable. Assume... and Representing the first and second elements in the real-time data matrix and historical data matrix, respectively. The and the first Given a vector of variables, the system performs a regularization transformation according to the following formula: Equation (1) Equation (2) In the formula, Represents the processed variable. Represents the original variable. and This is an independently generated and added Gaussian noise term. To ensure that the introduced noise does not mask the fault signals of the process itself, this embodiment introduces a relative coefficient to control the noise intensity. (For example, a value of 0.05) makes the variance of the noise proportional to the variance of the original variable. Specifically, the noise Follows a mean of 0 and a variance of The normal distribution, i.e. as well as ,in This corresponds to the variance of the original variables. In this way, this step effectively ensures that the correlation matrix (such as...) is consistent in subsequent orthogonal projection operations based on least squares. The process is reversible, thus ensuring the numerical stability and robustness of the entire progressive decoupling algorithm.

[0039] Regarding the processing details of the progressive temporal decoupling module, in some examples of embodiments of this application, linear temporal decoupling is performed. Orthogonal projection operations are used to separate the original data into time-independent linear components and historical data residuals that can be predicted from historical information. Hierarchical recursive nonlinear temporal decoupling is performed by using a parameterized activation function to perform nonlinear mapping on the output of the previous layer, and then orthogonally projecting the mapped result with the historical data residuals again to extract nonlinear feature residuals layer by layer. Data standardization is performed on the finally extracted deep time-independent components to convert them into standardized data with a mean of zero and a standard deviation of one, which serves as the input to the progressive feature representation module.

[0040] Here, the progressive temporal decoupling module completely removes the dynamic trends contained in historical data from the current data through a recursive approach that alternates between linear and nonlinear methods.

[0041] More specifically, linear temporal decoupling (i.e., layer 1 decoupling) is performed first. The system is based on a pre-constructed historical data matrix. Calculate the orthogonal projection operator and use it to project the real-time data matrix. The projection is onto the linear orthogonal complement space of the historical data space. The mathematical logic of this process follows the least squares principle, aiming to separate the dynamic components (i.e., predicted values) that can be linearly predicted by historical data from the static components (i.e., residuals) that cannot be predicted.

[0042] The specific orthogonal projection calculation is as follows: Equation (3) in, This represents the time-independent linear component of the first-layer output, effectively eliminating the most significant linear autocorrelation in the data.

[0043] Subsequently, the module enters the hierarchical recursive nonlinear temporal decoupling stage. To capture the deep nonlinear dynamic characteristics of the chemical process, the system will decouple the output from the previous layer. The input is fed into a non-linear activation function for feature mapping.

[0044] For example, a parameterized Swish activation function can be used. Its expression is ,in With adjustable parameters (e.g., 0.01), such functions possess smooth and non-monotonic properties, effectively preserving negative value information and uncovering hidden features. Mapped data New features deemed to contain potential temporal correlations are again subjected to orthogonal projection operations with historical data (or its mapping form). This process is repeated recursively. Each layer will filter out the dynamic correlations exposed by nonlinear transformation layer by layer, ultimately obtaining pure depth-time-independent components. .

[0045] Finally, to eliminate dimensional drift caused by multi-layer computation and meet the input requirements of subsequent networks, deep stationary data standardization is performed on the finally extracted components. System computation mean vector and standard deviation vector And perform a standardization transformation: Equation (4) After this processing, the output data is converted into a standard normal distribution with a mean of 0 and a standard deviation of 1, realizing the transformation from non-stationary data to stationary data. It also provides high-quality input with uniform scale for the subsequent progressive feature representation module, ensuring the sensitivity of the fault detection model to small changes.

[0046] In some examples of embodiments of this application, the progressive feature representation module employs a stacked neural component analysis network structure. More specifically, deep temporally irrelevant components are input into the first layer of the stacked neural component analysis network. In each layer, initial latent variables are extracted based on the input data of that layer, the variance of the initial latent variables is calculated and sorted in descending order, the number of major latent variables is determined based on the cumulative variance contribution rate, and the initial latent variables are then filtered to obtain the final latent variables of that layer. The reconstructed output of that layer is calculated based on the final latent variables, and the final latent variables are used as the input of the next layer, obtaining deep feature representations through hierarchical recursion.

[0047] In this embodiment, the core of the progressive feature representation module lies in utilizing a Stacked Neural Component Analysis (SNCA) structure to further extract statistically independent nonlinear features from standardized deep temporally uncorrelated components. Specifically, in each layer of the network (e.g., the first...),... In the first layer, the input data is processed through a weight matrix mapping and a sigmoid activation function to obtain a set of initial latent variables. To remove feature redundancy and retain key information, the system calculates the variance of all initial latent variables and arranges them in descending order. Subsequently, the cumulative variance contribution rate is calculated. ,in For the first The variance of each latent variable. When this ratio exceeds a preset threshold (e.g., 85%), the corresponding... The value represents the number of key latent variables retained at that level. Utilizing the principal component method in statistics, this ensures that subsequent processing focuses only on the critical variables that carry the majority of process variation information, thereby reducing noise interference.

[0048] After determining the number of primary latent variables, the module further performs refined latent variable screening and recursive propagation. Calculate the... The and the first The average variance of each latent variable is used as the standard deviation threshold, and a screening scoring mechanism based on the ReLU function is applied: Equation (5) This mechanism uses nonlinear truncation to reset or suppress the weights of minor latent variables with variances significantly below a threshold, thus obtaining the final, uncorrelated latent variables for that layer. Subsequently, the system uses the inverse matrix of the weights to calculate the reconstruction output of that layer (used for subsequent calculation of reconstruction error statistics), and directly uses the filtered final latent variables as input to the next layer.

[0049] Through a hierarchical recursive "extraction-filtering-transmission" mechanism, the network can abstract layer by layer, and finally output a deep feature representation that eliminates temporal correlation and satisfies statistical independence, which greatly improves the model's ability to represent complex nonlinear faults.

[0050] Regarding the construction of the progressive feature representation module, it can first employ Neural Component Analysis (NCA) to extract uncorrelated latent variables in the hidden layers, enabling the network to adaptively learn the nonlinear characteristics of chemical process data. Subsequently, a stacked NCA (SNCA) structure is used for multi-layer cascading, with each layer taking the latent variables extracted from the previous layer as input. Based on inheriting the abstract features of the upper layers, the original information is reintroduced for deeper independent component extraction. This process unfolds recursively in a hierarchical manner, after... q Layer stacking processing gradually extracts basic nonlinear components from shallow layers to highly abstract independent components from deep layers, ultimately obtaining deep feature representations that simultaneously satisfy temporal inconsistency and statistical independence, providing high-quality feature inputs for the detection of minor faults.

[0051] Regarding the implementation details of calculating online fault monitoring statistics, in some examples of embodiments of this application, a feature space correlation statistic is calculated. This is done by analyzing the correlation between the latent variables output in the last layer of the progressive feature representation module and the current sample and a preset number of samples in the past, to quantify the degree of abnormality in the variable relationships in the feature space. Furthermore, a residual space prediction error statistic is calculated. This is done by calculating the Euclidean distance between the input data and the reconstructed output of each layer of the progressive feature representation module, summing the Euclidean distances of all layers, and taking the average, to quantify the degree of process abnormality in the residual space.

[0052] More specifically, the calculation of fault monitoring statistics aims to comprehensively assess the operating status of a chemical process from two dimensions: the "characteristic space" and the "residual space." First, for the characteristic space correlation statistics (denoted as...),... Its core logic lies in monitoring whether deep latent variables deviate from the statistical distribution under normal operating conditions. Specifically, the system utilizes the last layer (the first layer) of the progressive feature representation module. Latent variable vector output by layer Combined with a preset sliding time window length Calculate the statistics: Equation (6) For the current moment and its past The latent variable magnitudes at each historical moment are cumulatively summed using squared values. Since the network training objective is to make the latent variables follow a specific independent distribution, when a failure occurs, the numerical magnitudes of these deep features or their temporal correlation patterns often shift significantly. By introducing a sliding window mechanism... Statistics can not only reflect the current abnormal magnitude, but also capture the cumulative effect of faults in a short period of time, thereby effectively improving the detection sensitivity of minor drift faults.

[0053] Secondly, regarding the residual space prediction error statistic (denoted as...) This metric measures the network model's ability to interpret the current input data. Unlike traditional methods that only calculate the reconstruction error of the final output, this embodiment uses an error averaging mechanism across all network layers. Specifically, the system traverses each layer of the progressive feature representation module, calculating the input data for that layer. Reconstructed data after network mapping and inverse mapping Calculate the squared Euclidean distance between them and take the average of the errors across all levels: Equation (7) In the formula, Representing the total number of network layers, it can simultaneously capture the disruption of local statistical features corresponding to shallow networks and the disruption of global nonlinear relationships corresponding to deep networks. If A significant increase in the value indicates that the variable correlation structure within the current data sample can no longer be reconstructed by a network trained on normal data, thus intuitively indicating process anomalies in the residual space.

[0054] In some examples of embodiments of this application, the pre-trained progressive temporal decoupling module is obtained through the following steps: constructing a training sample set based on historical data under normal operating conditions; inputting the training sample set into the progressive temporal decoupling module, performing multi-level recursive orthogonal projection and nonlinear mapping, obtaining the depth-time uncorrelated components, and calculating the mean and standard deviation of the depth-time uncorrelated components as standardization parameters to complete the configuration of the progressive temporal decoupling module.

[0055] In this embodiment, the pre-training or configuration process of the progressive temporal decoupling module mainly involves establishing a baseline for data distribution during the offline phase. First, the system constructs a training sample set based on historical data under normal operating conditions and inputs this dataset into the TOSA (Temporal Orthonormal Subspace Analysis) module. The module performs multi-level recursive operations on the training set according to the orthogonal projection principle (as shown in Equation 3). In each level, the orthogonal projection matrix is ​​calculated using the current batch of training data, extracting dynamic components that can be predicted from historical information and retaining unpredictable residuals. Through this process, the system is essentially learning the dynamic temporal structure of the chemical process under normal conditions, ultimately obtaining a set of deep time-independent components that represent the inherent nonlinear properties of the system. This results in a pure feature basis that has been freed from temporal autocorrelation.

[0056] Subsequently, to address the impact of the non-stationarity and dimensional differences in chemical engineering data on neural network training, the system performs parameterization on the aforementioned depth-time-independent components. Specifically, it calculates the mean vector of all training samples along the feature dimension. and standard deviation vector These two statistics constitute the standardized parameters of the module, following equation (4). These are the original depth-independent components. This is the standardized output. At this point, the module configuration is complete, and these two parameters ( and The data will be solidified and stored. In the subsequent online monitoring phase, regardless of how the distribution of real-time data fluctuates, the system will forcibly use this set of offline-obtained parameters to standardize the online decoupling results. This not only ensures that the data input to the subsequent SNCA network always maintains a uniform scale of zero mean and unit variance, but also effectively prevents the non-stationary drift of online data from masking the true fault characteristics, thus ensuring the consistency of model detection.

[0057] In some implementations, current and historical data are projected into mutually orthogonal subspaces to separate predictable dynamic components from unpredictable static linear features. Since orthogonal projection involves solving a least-squares problem, the historical data matrix needs to satisfy a full-rank condition to guarantee the existence of an inverse matrix. Therefore, a full-rank determination is performed on the historical data matrix before decoupling. If the condition is not met, adaptive regularization is used, adding independent Gaussian noise of controllable intensity to the data to ensure matrix invertibility and numerical stability. Subsequently, the obtained static features are mapped nonlinearly to obtain static feature residuals, while historical data are also mapped nonlinearly... The mapping yields historical data residuals (i.e., the nonlinear part that is useless for linear prediction). These two residuals are then re-inputted into the TOSA model for further decoupling analysis. This process unfolds recursively in a hierarchical manner, gradually eliminating temporal correlations in the data through l-level cascaded processing, ultimately obtaining a highly stationary deep feature representation. To eliminate differences in dimensions and numerical ranges between different data sources, the deep features and the original data are standardized, and their respective mean vectors and standard deviation vectors are calculated. The data is then converted into a standardized form with a mean of 0 and a standard deviation of 1, providing input data of a uniform scale for the subsequent progressive feature representation module.

[0058] In some examples of embodiments of this application, the pre-trained progressive feature representation module is obtained through the following steps: standardizing the depth-time-independent components using standardized parameters as input data for the progressive feature representation module; constructing an explicit loss function, which includes a Pearson correlation coefficient loss term to measure the uncorrelation between latent variables and a reconstruction error loss term to measure the difference between the network input and the reconstructed output; and using gradient descent to iteratively update the network parameters of the progressive feature representation module based on the explicit loss function until the network converges to obtain the trained progressive feature representation module.

[0059] In this embodiment, the training process of the progressive feature representation module aims to ensure that the latent variables extracted by the network can accurately reconstruct the original information while minimizing the statistical correlation between variables. First, the depth-temporally uncorrelated components are standardized using the standardized parameters (mean and standard deviation) obtained in the preceding steps, and these are used as input to the SNCA network, ensuring the stability of the network input data distribution. Subsequently, an explicit composite loss function is constructed. The function consists of two parts: the first part is the Pearson correlation coefficient loss term, which is designed to measure the lack of correlation between latent variables. This is achieved by calculating the correlation between any two different latent variables in the same stratum. and The sum of squares of the Pearson correlation coefficients between them, i.e.: Equation (8) In the formula, Describing covariance, Variance is represented by this term. Minimizing this term forces the network to decouple data features into statistically independent components, thereby eliminating redundancy.

[0060] The second part is the reconstruction error loss term. The mean squared error (MSE) is typically used to measure the difference between the network input and the reconstructed output after the encoding-decoding process, ensuring that the extracted features retain the key information of the original data. The final total loss function is a weighted sum of the two: Equation (9) By adjusting the weights Balancing the degree of decoupling with the accuracy of reconstruction.

[0061] Based on the loss function constructed above, the system uses gradient descent to iteratively optimize the network parameters. In practice, unlike traditional neural networks that rely solely on backpropagation for numerical approximation, the network weights are derived based on the objective function. and bias The analytical gradient formula makes the gradient calculation more accurate and computationally efficient.

[0062] In each iteration, the system updates the parameters of all levels based on the calculated gradient direction, as follows: Equation (10) In the formula, The learning rate is used. This iterative process continues until the loss function converges or the preset number of iterations is reached. At this point, the network parameters are fixed, and the finally trained progressive feature representation module can output deep features that simultaneously satisfy temporal inconsistency and statistical independence, providing a highly sensitive discrimination criterion for fault detection.

[0063] Here, the training and parameter optimization of the SNCA network are discussed. To achieve effective training of the progressive feature representation module, reconstruction error is introduced as the second loss component based on the constraint of mutual incorrelation, and the two losses are weighted and summed to construct the total loss function. A gradient descent-based parameter update strategy is used to optimize the network weights and biases.

[0064] In some examples of embodiments of this application, the preset control limits are obtained through the following steps: processing the training sample set using the configured progressive temporal decoupling module and the trained progressive feature representation module, and calculating the feature space correlation statistic and the residual space prediction error statistic corresponding to each training sample; fitting the probability density functions of the feature space correlation statistic and the residual space prediction error statistic respectively using the kernel density estimation method; and determining the control limits corresponding to the feature space correlation statistic and the residual space prediction error statistic respectively based on the given confidence level and the probability density functions.

[0065] Here, the determination of control limits is a crucial step connecting offline training and online monitoring. Specifically, after the progressive temporal decoupling module is configured and the progressive feature representation module training converges, the system re-inputs the training sample set, which represents the normal operating state of the process, into the aforementioned network model. For each sample in the training set, the system calculates the correlation of latent variables in the feature space according to the aforementioned statistical calculation logic. Reconstruction error of residual space Two sets of statistical sequences describing the fluctuation range under normal operating conditions were obtained.

[0066] In the calculation of control limits for fault detection statistics, the calculation A statistic, calculated by examining the sample and its past... T The second method monitors whether the correlation between latent variables in the last layer of the NCA network in each sample is abnormal, indicating whether the variable relationships in the feature space are abnormal; Statistics are used to monitor process anomalies in the residual space. The squared Euclidean distance between the inputs of all layers of the network and the reconstructed output is calculated and divided by the number of network layers. The control limits of each statistic are determined by the kernel density estimation method. The obtained control limits are used as the boundary threshold between normal operating conditions and fault conditions, and are used for fault judgment in the online detection stage.

[0067] To reconstruct a statistically "normal" baseline distribution, given that chemical process data often exhibit complex non-Gaussian distribution characteristics, traditional threshold determination methods based on the Gaussian assumption (such as...) are insufficient. The criteria often lead to a high false alarm rate or false negative rate, so this embodiment does not directly use the mean and variance to set the threshold.

[0068] Based on the obtained statistical series, the kernel density estimation (KDE) method is used to nonparametrically fit the probability density function (PDF) of each statistic. KDE smoothly estimates the overall distribution shape by stacking a kernel function (such as a Gaussian kernel) at each data point, accurately describing the skewness and multimodal characteristics of the data. After obtaining the probability density function... Then, the system determines the confidence level based on the given confidence level. (For example, targeting) The statistic is set at 99%, targeting The control limits were determined by setting the statistic to 85%. .

[0069] Solve the integral equations that satisfy the given conditions threshold Therefore, under normal operating conditions, the value of the statistic is: The probability falls within this control limit. Control limits determined in this way possess both statistical rigor and adaptability to the complex random fluctuations of chemical processes, thus effectively ensuring the accuracy of alarms during online detection.

[0070] Figure 2 A flowchart illustrating another example of a chemical process fault detection method based on a progressively decoupled representation learning network according to an embodiment of this application is shown.

[0071] First, data collection and data matrix construction are carried out.

[0072] For example, the effectiveness of the proposed method was verified using the Tennessee Eastman (TE) chemical process simulation dataset. This chemical process simulation platform comprises five main units: reactor, condenser, gas-liquid separator, circulating compressor, and stripping tower, involving four reactants, two products, one inert substance, and one by-product. It exhibits highly nonlinear, strongly coupled, and multivariable characteristics. This invention selected 33 key process variables, including parameters such as temperature, pressure, flow rate, liquid level, and composition. After data acquisition, a real-time data matrix was constructed according to the steps described in S1 above. and historical data matrix The number of variables in this embodiment Historical time series length Number of variables in historical data matrix .

[0073] Then, a progressive timing decoupling module is constructed.

[0074] Specifically, the historical data matrix is ​​used for row full-rank determination. This is to ensure the subsequent time-series decoupling process. This item does indeed have an inverse matrix, and the historical data matrix needs to be determined beforehand. Is it the full rank? If If the rows are full, then it is a square matrix. It is reversible. If If the rows are not full rank, then perform adaptive regularization: Equation (11) in, ; ,Right now and Represent and The first in and One variable; and It was added manually. and Independent Gaussian noise: Equation (12) In the formula, and The relative coefficient representing the control noise intensity is set to 0.05 in this embodiment, which is equivalent to 5% of the standard deviation of the original variable. The mean is 0 and the variance is . The normal distribution; The mean is 0 and the variance is . It follows a normal distribution.

[0075] Linear temporal decoupling is performed (layer 1). Specifically, orthogonal projection techniques are used to calculate the time-independent linear components in the original data. and residual information from historical data that did not contribute to linear prediction : Equation (13) No. Layer-time decoupling. The static features of the previous layer are processed using a parameterized Swish activation function. After mapping, perform a new orthogonal projection with historical data: Equation (14) In the formula, , Activation function It can be obtained from the following formula: Equation (15) In the formula, It is a very small, fixed constant, and its value can be 0.01; Perform depth-stationary data standardization. Standardize the extracted depth-temporally uncorrelated components as shown in the following formula: Equation (16) In the formula, It is standardized data; and Represent The mean and standard deviation (of the feature dimensions); similarly, this step ultimately yields... n OK r Column data matrix This can also be called the deep temporally uncorrelated component, which will serve as the first layer input to the subsequent progressive feature representation module.

[0076] Furthermore, a progressive feature representation module is constructed.

[0077] For the first layer feature representation, the extracted depth-temporally uncorrelated components are... The input to the hidden layer is obtained by applying weights and biases. The latent variables of the first layer network are obtained after passing through the sigmoid activation function. ;in, , .

[0078] calculate variance And arrange them in descending order.

[0079] Calculate the cumulative variance contribution rate : Equation (17) in It is the number of variables selected as the primary latent variables.

[0080] Calculate the standard deviation threshold .

[0081] Calculate the screening score for each latent variable. .

[0082] Equation (18) in, It is a very small constant to prevent numerical instability caused by a denominator term of 0, for example. .

[0083] Calculate hidden layer output The reconstructed values ​​of the first layer network are obtained by inverting the weights and the negative of the biases. .

[0084] No. Layer feature representation, Latent variables of the layer Sent to the Layer. Then repeat the steps of the first layer to obtain the [number]th layer. Latent variables of the layer and reconstructed values In the formula .

[0085] Furthermore, the training and parameter optimization of the SNCA network are performed.

[0086] For each layer of latent variables Construct the objective function And minimize it to ensure that the latent variables are uncorrelated, thereby achieving the effect of removing feature redundancy: Equation (19) in, For the sake of clarity in the following steps, the numerator term will be abbreviated as... .

[0087] Solving for weights w and bias b The gradient is shown in the following equation: Equation (20) in: Equation (21) In the above formula Represents the sample mean. , and .

[0088] Update weights using gradient descent. w and bias b As shown in the following formula: Equation (22) In the formula, This represents the number of iterations in the training process. This represents the learning rate. In this embodiment... , .

[0089] Then, calculate the fault detection statistics. Control limits.

[0090] Specifically, for the first Samples at time ( ), calculate the sample and its past T The first sample Latent variables of layered NCA networks and The correlation between them, as Statistic: Equation (23) in, , and The control limit for this indicator is generated by kernel probability density estimation, and in this embodiment, the confidence level is set to 99%.

[0091] For the t For a sample at time step 1, calculate the average of the sum of squares of the reconstruction errors of that sample across all layers in the SNCA network, and then calculate the squared prediction error. SPE (Squared Prediction Error) statistic: Equation (24) The control limit for this indicator is generated by kernel probability density estimation, for example, with a confidence level of 85%.

[0092] Finally, perform online fault detection.

[0093] For real-time data, the statistical value of the data is calculated based on statistical indicators. If this value exceeds the calculated control limit, it indicates that a fault has occurred. The test data for the TE process starts from the 161st sample ( The fault was introduced starting with the 960th sample.

[0094] Figure 3 A schematic diagram of the deep neural network used in the progressive timing decoupling module and the progressive feature representation module according to embodiments of this application is shown.

[0095] like Figure 3 As shown on the left, the progressive timing decoupling module adopts a dual-stream input architecture, where the upper branch inputs real-time data ( The branch below allows you to input historical data. This module performs recursive decoupling through a hierarchical structure: the first layer separates the initial dynamic components through orthogonal projection. ) and historical residuals ( Then, a nonlinear activation function was used. (as shown in the picture) and The output of the previous layer is mapped to a high-dimensional feature space, and orthogonal projection is performed again. This process... The process is recursively performed at each level, peeling away nonlinear temporal correlations layer by layer, ultimately outputting deeply stationary variables. This refers to the depth-time uncorrelated component.

[0096] like Figure 3 As shown on the right, the progressive feature representation module receives standardized stationary variables and employs a cascaded structure of stacked neural component analysis (SNCA). This module consists of... It is composed of several sub-networks connected in series, each layer (based on the first...) For example, see the layer. Figure 3 (Enlarged illustration in the bottom right corner) All processes involve three stages: encoding, filtering, and decoding. Input data First, weights are mapped to the hidden space. In order to extract the initial latent variables Subsequently, the system calculates the variance of each latent variable. And based on the threshold Nonlinear screening was performed to retain the main latent variables that contributed significantly to variance. Redundant features were removed. The final latent variables after screening. On the one hand, it is used to reconstruct the output. On the one hand, it is used to calculate the reconstruction error, and on the other hand, it is directly used as the input to the next layer of the network (i.e. This allows for the gradual abstraction and refinement of features from shallow to deep independent features.

[0097] The table below shows the descriptions of 20 faults and the difference between the detection rate and false alarm rate of the proposed method for these 20 faults in the embodiments of this application to verify the performance of the method. It should be noted that... I 2 and SPEIf either of the two indicators detects a fault, it means that a fault has occurred.

[0098] Figures 4-23 The detection results of IDV (1) to IDV (20) faults in the Tennessee Eastman process according to the embodiments of this application are shown respectively.

[0099] Table 1. Descriptions and detection results of 20 faults.

[0100] In the figure, Process Monitoring Results represent process monitoring results, Sampling Time represents the sampling time (horizontal axis), Fault Time represents the time of fault occurrence, Fault Operation represents the fault condition, Threshold represents the control limit, and Normal Operation represents the normal condition.

[0101] The following is an explanation of the unknown faults mentioned in the table above: Table 2 Description of Unknown Faults

[0102] Among them, the fault amplitude quantifies the degree of deviation of process variables from the normal operating state after an unknown fault is injected into the TE process, that is, the severity of the fault.

[0103] The chemical reaction formulas involved in the TE process are as follows: Equation (25) As can be seen from the above formula, the TE process has four gaseous reactants A, C, D and E, as well as an inert component B. The reactants and inert component are fed into the reactor, and the products are G and H.

[0104] Here, we present the detection results of the algorithm invented in this patent for 20 faults (corresponding to Table 2). In the figure, the blue solid line represents the value of the statistical quantity of normal data, and the red solid line represents the value of the statistical quantity after the fault is introduced. The black dashed line represents the value of the control limit, and the orange dashed line represents the time of fault introduction.

[0105] It is important to note that I 2 If either the SPE or MIRV indicator detects a fault, it indicates that a fault has occurred. (Refer to Table 2 and...) Figures 4-23 The following conclusions can be drawn from the content: 1) This method I 2 and SPEBoth indicators demonstrate significant advantages in detecting faults 1, 2, 5, 6, 7, and 14. The statistical values ​​immediately exceed their corresponding thresholds early in the fault occurrence, and the subsequent values ​​continue to climb, far exceeding the thresholds. It is worth noting that fault 5 is more difficult to detect. This is because fault 5 represents a sudden change in the cooling water inlet temperature, occurring at the 161st sampling point. The change in cooling water temperature causes other variables to deviate from normal values ​​during the process. However, due to the actions of various loop controllers, most of these variables are gradually brought back to normal values ​​after approximately the 200th sampling point, making the entire process appear to have returned to normal. It is crucial to note that the cooling water temperature deviation persists and does not return to normal values ​​despite the recovery of other variables. However, the statistical values ​​of the two indicators in this method do not decrease due to this. This indicates that the algorithm can penetrate the compensation effect of the control loop and directly detect residual faults that have not been eliminated.

[0106] 2) This method performs exceptionally well for difficult-to-detect nonlinear faults. In the TE process, faults 3, 9, and 15 are the most difficult to detect because, after these faults occur, the state of the observed variables is almost identical to that under normal conditions. This method… I 2 The indicators quickly exceeded the threshold after faults 3 and 9 occurred and remained stable, effectively triggering the alarm. For fault 15, although the indicator did not exceed the threshold in the early stages, its trajectory had clearly deviated from the normal range, and the fault was finally successfully captured after the 300th sampling point.

[0107] 3) Leveraging the cumulative effect of the sliding window mechanism, this method... I 2 Indicator Ratio SPE The indicators are more sensitive and can effectively detect the vast majority of faults. Combined with the fusion decision-making strategy of "alarm upon any indicator exceeding the limit," this method relies on… I 2 The high sensitivity of the indicators enabled successful detection of most faults. As can be seen from the monitoring curves for faults 4, 8, 10-13, and 16-20, I 2 The statistic value of the indicator immediately exceeded its corresponding threshold in the early stages of the failure, and the value continued to climb thereafter.

[0108] The fault detection method for chemical processes proposed in this invention employs a progressive TOSA multi-layer cascade, which can extract features layer by layer from shallow time-series patterns to deep, complex nonlinear relationships, effectively capturing subtle changes in faults. Addressing the non-stationary nature of chemical process data, which exhibits time-varying mean and variance, the feature extraction process eliminates the need for preprocessing standardization, directly removing the time-series autocorrelation of non-stationary data. This ultimately yields a highly stationary feature representation with mean and variance that do not change over time, achieving an effective conversion from non-stationary to stationary data and providing a stable and reliable data foundation for fault detection.

[0109] In the progressive temporal decoupling module of the fault detection method for chemical processes, a lightweight network design is adopted, which is computationally efficient and easy to deploy in industry. Unlike traditional methods that rely on deep neural networks with large numbers of parameters for temporal modeling, the TOSA module in this embodiment is based on matrix orthogonal projection operations, and its network structure is essentially parameterless. This module directly decouples temporal correlations through a series of efficient numerical calculations, avoiding the huge computational overhead and overfitting risk caused by training millions of weight parameters. This lightweight design makes it require very little hardware computing power, enabling it to run efficiently on edge computing devices or low-configuration industrial control computers in industrial settings, meeting the stringent real-time requirements of chemical process fault detection tasks.

[0110] In the progressive feature characterization module of the fault detection method for chemical processes, a layer-by-layer abstraction feature extraction mechanism is adopted. The shallow network captures the local nonlinear patterns and basic feature transformations of the chemical process, while the deep network, based on the shallow features, performs higher-level synthesis and abstraction to form a deep representation of complex nonlinear relationships. This hierarchical structure can effectively characterize the multi-scale and multi-level complex patterns that are common in chemical processes, breaking through the limitations of the expressive power of similar methods.

[0111] In fault detection methods for chemical processes, compared to traditional mechanism-based modeling methods that require in-depth understanding of process mechanisms, establishment of complex mathematical models, and tedious parameter identification, this method only needs to collect historical data under normal operating conditions to complete model training, significantly shortening the cycle from method development to actual deployment. Furthermore, the method provided in this application has a flexible parameter adjustment mechanism, capable of adaptive configuration according to the characteristics of different chemical processes, exhibiting excellent adaptability to the complex data characteristics commonly found in chemical processes, such as high dimensionality, strong nonlinearity, and dynamic time-varying nature. This plug-and-play characteristic makes it easy to quickly deploy and promote its application in actual industrial plants, eliminating the need for extensive customized development for each specific process, thus demonstrating good engineering practicality and economic benefits.

[0112] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of combined actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Secondly, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application. In the above embodiments, the descriptions of each embodiment have their own emphasis; for parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0113] In some embodiments, this application also provides a computer program product, the computer program product including a computer program stored on a non-volatile computer-readable storage medium, the computer program including program instructions, which, when executed by a computer, cause the computer to execute any of the above-described chemical process fault detection methods based on progressive decoupling characterization learning networks.

[0114] In some embodiments, this application also provides an electronic device comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to execute a chemical process fault detection method based on a progressively decoupled characterization learning network.

[0115] The apparatus described in the embodiments of this application can be used to execute the chemical process fault detection method based on a progressively decoupled characterization learning network according to the embodiments of this application, and accordingly achieve the technical effects achieved by the chemical process fault detection method based on a progressively decoupled characterization learning network according to the embodiments of this application, which will not be elaborated further here. In the embodiments of this application, the relevant functional modules can be implemented using a hardware processor.

[0116] Figure 24 This is a schematic diagram of the hardware structure of an electronic device for implementing a chemical process fault detection method based on a progressively decoupled representation learning network, as provided in another embodiment of this application. Figure 24 As shown, the device includes: One or more processors 2410 and memory 2420, Figure 24 Take the 2410 processor as an example.

[0117] The device for implementing the chemical process fault detection method based on the progressive decoupled characterization learning network may further include: an input device 2430 and an output device 2440.

[0118] The processor 2410, memory 2420, input device 2430, and output device 2440 can be connected via a bus or other means. Figure 24 Taking the example of a connection between China and Israel via a bus.

[0119] The memory 2420, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the chemical process fault detection method based on a progressively decoupled characterization learning network in the embodiments of this application. The processor 2410 executes various server functions and data processing by running the non-volatile software programs, instructions, and modules stored in the memory 2420, thereby implementing the chemical process fault detection method based on a progressively decoupled characterization learning network described in the above embodiments.

[0120] The memory 2420 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the device, etc. Furthermore, the memory 2420 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 2420 may optionally include memory remotely located relative to the processor 2410, and these remote memories may be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0121] Input device 2430 can receive input digital or character information and generate signals related to user settings and function control of the device. Output device 2440 may include display devices such as a display screen.

[0122] The one or more modules are stored in the memory 2420. When executed by the one or more processors 2410, they execute the chemical process fault detection method based on progressive decoupling representation learning network in any of the above method embodiments.

[0123] The above-described product can perform the methods provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects for performing the methods. Technical details not described in detail in this embodiment can be found in the methods provided in the embodiments of this application.

[0124] The electronic devices in this application embodiments exist in various forms, including but not limited to: (1) Mobile communication devices: These devices are characterized by their mobile communication capabilities and primarily aim to provide voice and data communication. These terminals include: smartphones (e.g., iPhones), multimedia phones, feature phones, and low-end phones, etc.

[0125] (2) Ultra-mobile personal computer devices: These devices fall under the category of personal computers, possessing computing and processing capabilities, and generally also have mobile internet access features. These terminals include PDAs, MIDs, and UMPCs, such as the iPad.

[0126] (3) Portable entertainment devices: These devices can display and play multimedia content. This category includes audio and video players (such as iPods), handheld game consoles, e-book readers, as well as smart toys and portable car navigation devices.

[0127] (4) Server: A device that provides computing services. The components of a server include a processor, hard disk, memory, system bus, etc. Servers are similar to general computer architectures, but because they need to provide highly reliable services, they have higher requirements in terms of processing power, stability, reliability, security, scalability, and manageability.

[0128] (5) Other electronic devices with data interaction functions.

[0129] In some embodiments, this application also provides a mobile platform on which the computer device described in any embodiment of this application is installed. The mobile platform includes, but is not limited to, vehicles, tracked robots, bipedal robots, quadrupedal robots, etc., wherein the vehicle can be a passenger car, pickup truck, truck, etc. It should be noted that the above are merely examples, and this application does not limit the specific form of the mobile platform.

[0130] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0131] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0132] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for fault detection in chemical processes based on a progressively decoupled representation learning network, comprising: Acquire real-time sampling data of the chemical process, and construct a real-time data matrix containing current time information and a historical data matrix containing information of a preset number of past times based on the real-time sampling data; The real-time data matrix and the historical data matrix are input into a pre-trained progressive temporal decoupling module. The progressive temporal decoupling module performs multi-level recursive orthogonal projection and nonlinear mapping on the input data to eliminate temporal correlation and output standardized deep temporally uncorrelated components. The deep time-independent components are input into a pre-trained progressive feature representation module. The progressive feature representation module is used to calculate the latent variables and reconstructed outputs of each layer of the network. Based on the latent variables and reconstructed outputs, the corresponding online fault monitoring statistics are calculated. If the online fault monitoring statistics exceed the preset control limit, the chemical process is determined to be in a fault state and an alarm is triggered.

2. The method according to claim 1, characterized in that, The construction of a real-time data matrix containing current time information and a historical data matrix containing information from a predetermined number of past times based on the real-time sampled data includes: Based on the preset historical time series length, for each real-time sampling moment, process variable data of a preset number of consecutive sampling moments before that sampling moment are obtained; The process variable data of the continuous preset number of sampling times are concatenated and spliced ​​in chronological order to form a historical information vector that can characterize the dynamic characteristics of the process. The real-time data matrix is ​​constructed based on the historical information vector at the current moment, and the historical data matrix is ​​constructed based on the historical information vector at a preset number of past moments.

3. The method according to claim 1, characterized in that, Before performing multi-level recursive orthogonal projection and nonlinear mapping on the input data using the progressive temporal decoupling module, the method further includes: Perform row full rank determination on the real-time data matrix and the historical data matrix; If the matrix is ​​not of full rank, the relative coefficient of the control noise intensity is calculated, and independent Gaussian noise is introduced into the real-time data matrix and the historical data matrix based on the relative coefficient to perform adaptive regularization processing, thereby ensuring the orthogonal projection invertibility of matrix operations.

4. The method according to claim 1, characterized in that, The step of using the progressive temporal decoupling module to perform multi-level recursive orthogonal projection and nonlinear mapping on the input data includes: Perform linear time series decoupling, and separate the original data into time-independent linear components and historical data residuals that can be predicted by historical information through orthogonal projection operations; The hierarchical recursive nonlinear temporal decoupling is performed by using a parameterized activation function to perform nonlinear mapping on the output of the previous layer, and then the mapped result is orthogonally projected with the historical data residual to extract nonlinear feature residuals layer by layer. The extracted depth-time-independent components are subjected to data standardization processing to convert them into standardized data with a mean of zero and a standard deviation of one, which is then used as input to the progressive feature representation module.

5. The method according to claim 1, characterized in that, The progressive feature representation module employs a stacked neural component analysis network structure. The calculation of latent variables and reconstructed outputs at each network level using this module includes: The depth-temporally irrelevant components are input into the first layer of the stacked neural component analysis network; In each layer of the network, initial latent variables are extracted based on the input data of that layer, the variance of the initial latent variables is calculated and sorted in descending order, the number of major latent variables is determined based on the cumulative variance contribution rate, and the initial latent variables are screened accordingly to obtain the final latent variables of that layer. The reconstructed output of this layer is calculated based on the final latent variables, and the final latent variables are used as the input of the next layer of the network to obtain deep feature representations through hierarchical recursion.

6. The method according to claim 1, characterized in that, The online fault monitoring statistics calculated based on the latent variables and the reconstructed output include: The correlation statistics of the feature space are calculated. By analyzing the correlation between the latent variables output by the current sample and a preset number of samples in the last layer of the progressive feature representation module, the degree of abnormality of the variable relationship in the feature space is quantified. The residual space prediction error statistic is calculated by calculating the Euclidean distance between the input data and the reconstructed output of each layer of the progressive feature representation module, and then summing the Euclidean distances of all layers and taking the average value to quantify the degree of process anomaly in the residual space.

7. The method according to claim 1, characterized in that, The pre-trained progressive temporal decoupling module is obtained through the following steps: A training sample set is constructed based on historical data under normal operating conditions; The training sample set is input into the progressive temporal decoupling module, which performs multi-level recursive orthogonal projection and nonlinear mapping to obtain depth-time uncorrelated components. The mean and standard deviation of the depth-time uncorrelated components are then calculated as standardization parameters to complete the configuration of the progressive temporal decoupling module.

8. The method according to claim 7, characterized in that, The pre-trained progressive feature representation module is obtained through the following steps: The depth-time-independent components are standardized using the standardized parameters and used as input data for the progressive feature representation module. Construct an explicit loss function, which includes a Pearson correlation coefficient loss term to measure the lack of correlation between latent variables and a reconstruction error loss term to measure the difference between network input and reconstruction output. The gradient descent method is used to iteratively update the network parameters of the progressive feature representation module based on the explicit loss function until the network converges, so as to obtain the trained progressive feature representation module.

9. The method according to claim 8, characterized in that, The preset control limits are obtained through the following steps: The training sample set is processed using the configured progressive temporal decoupling module and the trained progressive feature representation module to calculate the feature space correlation statistic and residual space prediction error statistic for each training sample. The kernel density estimation method is used to fit the probability density functions of the correlation statistic of the feature space and the prediction error statistic of the residual space, respectively. Based on the given confidence level, the control limits corresponding to the feature space correlation statistic and the control limits corresponding to the residual space prediction error statistic are determined according to the probability density function.

10. A computer device, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1-9.