Method and device for early warning of faults in chemical processes across operating conditions
Patent Information
- Application Number
- CN202610628449.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-08
- Publication Date
- 2026-09-04
AI Technical Summary
经过研究发现,相关技术中的化工过程中的多步预测故障预警任务存在以下主要不足,传统故障检测与诊断方法存在明显滞后性,目前的数据驱动预测模型缺乏跨工况泛化能力,故障预警下的跨工况泛化面临更严峻挑战
Smart Images

Figure CN122694014A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of chemical process safety monitoring technology, and in particular to a method and device for cross-condition fault early warning of chemical processes based on hierarchical domain invariant alignment networks. Background Technology
[0002] With the continuous expansion of chemical plant scale and the increasing level of system integration, abnormal operating conditions and failure events may lead to significant economic losses, environmental risks, and even catastrophic accidents. Data-driven process monitoring and fault early warning have become key technical means to ensure the safety of modern chemical processes and product quality. Research has found that multi-step predictive fault early warning tasks in chemical processes have the following main shortcomings: traditional fault detection and diagnosis methods have significant lag; current data-driven prediction models lack cross-condition generalization capabilities; and cross-condition generalization under fault early warning faces even more severe challenges. Summary of the Invention
[0003] In view of this, this disclosure proposes a method and device for cross-condition fault early warning of chemical processes, which can realize domain-invariant learning of chemical process prediction models under multiple conditions, so that they can be directly deployed without any retraining on completely unseen conditions.
[0004] According to one aspect of this disclosure, a method for cross-condition fault early warning in chemical processes is provided. The method includes: acquiring a first process data stream of a target process flow under a target condition, the first process data stream including observations of multiple first process variables at multiple first historical moments; inputting the preprocessed first process data stream into a trained prediction model to obtain a prediction result, the prediction result including predicted values of the multiple first process variables at multiple future moments; determining an early warning scheme for the target condition based on the prediction result and preset early warning conditions; wherein, the method further includes a training process of the prediction model, the training process including the following steps: constructing a training sample set based on second process data streams of sample process flows under multiple sample conditions, the second process data stream including observations of each second process variable at multiple second historical moments, each sample condition being different from the target condition; training an initialized prediction model based on the training sample set to obtain a trained prediction model.
[0005] In one possible implementation, constructing a training sample set based on the second process data stream of sample process flows under multiple sample operating conditions includes: calculating the mean and standard deviation of the second process variable based on all observations corresponding to each second process variable; standardizing each observation of the second process variable based on the mean and standard deviation of each second process variable; for each second process variable, using the standardized observations at a first number of second historical times before the target time as historical samples, and using the standardized observations at a second number of second historical times after the target time as prediction samples, to construct the training sample set, which includes multiple sample pairs, each sample pair including historical samples and prediction samples for the same target time, wherein the number of target times is preset to be multiple.
[0006] In one possible implementation, the prediction model includes an encoder comprising a plurality of self-attention layers connected in sequence, each self-attention layer being connected to a residual domain adaptation module; the self-attention layers are used to extract features from the input to obtain hierarchical features, the input of the self-attention layer including the input embedding sequence or the hierarchical features output by the previous self-attention layer; the residual domain adaptation module is used to transform the hierarchical features output by the self-attention layers to which it is connected into sample-level aligned features.
[0007] In one possible implementation, the method further includes: performing a linear transformation on the historical input matrix corresponding to all historical samples in the training sample set to obtain a first matrix; and performing position encoding on the first matrix to obtain the input embedding sequence.
[0008] In one possible implementation, training the initialized prediction model based on the training sample set includes: calculating the hierarchical multi-kernel maximum mean difference value based on the sample-level alignment features output by all residual domain adaptation modules; and calculating the mean squared error value based on the sample-level alignment features output by all residual domain adaptation modules and the prediction samples in the training sample set; using the hierarchical multi-kernel maximum mean difference value and the mean squared error value as objective function values, and minimizing the objective function value as the training objective, adjusting the parameters of the encoder.
[0009] In one possible implementation, adjusting the parameters of the encoder includes: updating the alignment weights of each self-attention layer based on the sample-level alignment features output by the residual domain adaptation module, wherein the alignment weights are used to adjust the alignment intensity of the self-attention layers.
[0010] In one possible implementation, the preprocessing of the first process data stream includes: obtaining standardized parameters corresponding to each first process variable, wherein the standardized parameters refer to the mean and standard deviation calculated based on the observations of the same second process variable at multiple second historical moments; standardizing each observation of the first process variable according to the standardized parameters corresponding to each first process variable; and for each first process variable, using the standardized observations at a first number of first historical moments as an input matrix, wherein the input matrix is used as the input of the trained prediction model.
[0011] In one possible implementation, the warning conditions include warning conditions for a target process variable among the plurality of first process variables; determining the warning scheme for the target operating condition based on the prediction result and the preset warning conditions includes: if the predicted value of the target process variable at the future time is less than a preset first threshold, or the predicted value of the target process variable at the future time is greater than a preset second threshold, then the warning scheme for the target operating condition includes activating a fault alarm for the target process variable, wherein the first threshold is lower than the second threshold.
[0012] According to another aspect of this disclosure, a cross-condition fault early warning device for chemical processes is provided. The device includes: a data acquisition module for acquiring a first process data stream of a target process flow under a target condition, the first process data stream including observations of multiple first process variables at multiple first historical moments; a fault prediction module for inputting the preprocessed first process data stream into a trained prediction model to obtain a prediction result, the prediction result including predicted values of the multiple first process variables at multiple future moments; and an early warning generation module for determining an early warning scheme for the target condition based on the prediction result and preset early warning conditions. The device further includes a model training module for executing a training process of the prediction model, the training process including: constructing a training sample set based on second process data streams of sample process flows under multiple sample conditions, the second process data stream including observations of each second process variable at multiple second historical moments, each sample condition being different from the target condition; and training an initialized prediction model based on the training sample set to obtain a trained prediction model.
[0013] In one possible implementation, constructing a training sample set based on the second process data stream of sample process flows under multiple sample operating conditions includes: calculating the mean and standard deviation of the second process variable based on all observations corresponding to each second process variable; standardizing each observation of the second process variable based on the mean and standard deviation of each second process variable; for each second process variable, using the standardized observations at a first number of second historical times before the target time as historical samples, and using the standardized observations at a second number of second historical times after the target time as prediction samples, to construct the training sample set, which includes multiple sample pairs, each sample pair including historical samples and prediction samples for the same target time, wherein the number of target times is preset to be multiple.
[0014] In one possible implementation, the prediction model includes an encoder comprising a plurality of self-attention layers connected in sequence, each self-attention layer being connected to a residual domain adaptation module; the self-attention layers are used to extract features from the input to obtain hierarchical features, the input of the self-attention layer including the input embedding sequence or the hierarchical features output by the previous self-attention layer; the residual domain adaptation module is used to transform the hierarchical features output by the self-attention layers to which it is connected into sample-level aligned features.
[0015] In one possible implementation, the apparatus further includes a sample processing module, configured to: perform a linear transformation on the historical input matrix corresponding to all historical samples of the training sample set to obtain a first matrix; and perform position encoding on the first matrix to obtain the input embedding sequence.
[0016] In one possible implementation, training the initialized prediction model based on the training sample set includes: calculating the hierarchical multi-kernel maximum mean difference value based on the sample-level alignment features output by all residual domain adaptation modules; and calculating the mean squared error value based on the sample-level alignment features output by all residual domain adaptation modules and the prediction samples in the training sample set; using the hierarchical multi-kernel maximum mean difference value and the mean squared error value as objective function values, and minimizing the objective function value as the training objective, adjusting the parameters of the encoder.
[0017] In one possible implementation, adjusting the parameters of the encoder includes: updating the alignment weights of each self-attention layer based on the sample-level alignment features output by the residual domain adaptation module, wherein the alignment weights are used to adjust the alignment intensity of the self-attention layers.
[0018] In one possible implementation, the preprocessing of the first process data stream includes: obtaining standardized parameters corresponding to each first process variable, wherein the standardized parameters refer to the mean and standard deviation calculated based on the observations of the same second process variable at multiple second historical moments; standardizing each observation of the first process variable according to the standardized parameters corresponding to each first process variable; and for each first process variable, using the standardized observations at a first number of first historical moments as an input matrix, wherein the input matrix is used as the input of the trained prediction model.
[0019] In one possible implementation, the warning conditions include warning conditions for a target process variable among the plurality of first process variables; determining the warning scheme for the target operating condition based on the prediction result and the preset warning conditions includes: if the predicted value of the target process variable at the future time is less than a preset first threshold, or the predicted value of the target process variable at the future time is greater than a preset second threshold, then the warning scheme for the target operating condition includes activating a fault alarm for the target process variable, wherein the first threshold is lower than the second threshold.
[0020] According to another aspect of this disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above-described method.
[0021] According to another aspect of this disclosure, a non-volatile computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the above-described method.
[0022] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the above-described method.
[0023] Other features and aspects of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description
[0024] The accompanying drawings, which are included in and form part of this specification, illustrate exemplary embodiments, features, and aspects of this disclosure together with the specification and serve to explain the principles of this disclosure.
[0025] Figure 1 The flowchart illustrates the online early warning stage in the fault early warning method provided in this embodiment.
[0026] Figure 2a This diagram illustrates the process variable prediction provided by an embodiment of the present disclosure.
[0027] Figure 2b This diagram illustrates the alarm decision logic provided in an embodiment of the present disclosure.
[0028] Figure 3 The flowchart illustrates the offline training phase of the fault warning method provided in this embodiment.
[0029] Figure 4 This diagram illustrates the multi-condition domain generalization provided in the embodiments of this disclosure.
[0030] Figure 5 This diagram illustrates the overall architecture of HIAN and dual-stream training provided in an embodiment of this disclosure.
[0031] Figure 6 A TEP flowchart provided in an embodiment of this disclosure is shown.
[0032] Figure 7 This illustrates a UMAP clustering visualization diagram of FCC operating conditions provided in an embodiment of this disclosure.
[0033] Figure 8 This diagram shows a comparison of the 5-step advance prediction results for FCC operating condition 6 without any observed conditions, provided in an embodiment of this disclosure.
[0034] Figure 9 A block diagram of a chemical process cross-condition fault early warning device provided in an embodiment of this disclosure is shown. Detailed Implementation
[0035] Various exemplary embodiments, features, and aspects of this disclosure will now be described in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.
[0036] As used herein, the terms “comprising,” “including,” “having,” or variations thereof are open-ended and include one or more of the stated features, integrals, elements, steps, components, or functions, but do not exclude the presence or addition of one or more other features, integrals, elements, steps, components, functions, or groups thereof.
[0037] When an element is referred to as “connected,” “coupled,” “responding,” or a variation thereof relative to another element, it may be directly connected, coupled, or responding to another element, or there may be an intermediate element present.
[0038] Although the terms first, second, third, etc., may be used herein to describe various elements / operations, these elements / operations should not be limited by these terms. These terms are only used to distinguish one element / operation from another. Therefore, without departing from the teachings of the inventive concept, a first element / operation in some embodiments may be referred to as a second element / operation in other embodiments.
[0039] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments.
[0040] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.
[0041] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, data stored, data displayed, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant regions.
[0042] To facilitate understanding of the technical solutions provided by the embodiments of this disclosure by those skilled in the art, the technical environment for implementing the technical solutions will be described below.
[0043] As chemical plants continue to expand in scale and their system integration levels continue to improve, abnormal operating conditions and malfunctions may lead to significant economic losses, environmental risks, and even catastrophic accidents. Data-driven process monitoring and fault early warning have become key technical means to ensure the safety of modern chemical processes and the quality of products.
[0044] Multi-step predictive fault early warning in chemical processes is a technical approach that uses real-time monitoring of equipment operating data and artificial intelligence (AI) algorithms to predict potential fault risks at multiple future time points, thereby enabling proactive maintenance measures. The core of this task lies in overcoming the limitations of traditional single-step prediction by establishing dynamic models to capture fault evolution patterns, providing full-cycle safety assurance for high-temperature, high-pressure, flammable, and explosive chemical environments. "Multi-step" refers to multiple future moments. Unlike single-step prediction, which can only predict the state at the next point in time, multi-step prediction can output the operating trend for several future time points.
[0045] Research has revealed the following major shortcomings in related technologies. Firstly, traditional fault detection and diagnosis methods exhibit significant lag. Traditional fault detection and diagnosis (FDD) typically triggers alarms only after a fault has occurred and abnormal symptoms are detectable, representing a passive response mode. For slowly evolving, latent faults—such as valve / actuator jamming, sensor drift, and catalyst coking—this passive paradigm suffers from a significant detection delay, leaving operators with insufficient response time and amplifying accident risks. Therefore, fault prognosis, as a proactive approach, is gaining increasing attention. Its goal is to achieve early prediction and warning of abnormal events by multi-steply predicting the future trajectories of key process variables before a fault triggers an alarm.
[0046] Secondly, current data-driven prediction models lack cross-condition generalization capabilities. In actual production, chemical processes frequently switch between multiple steady-state conditions due to factors such as changes in raw material composition, product specification switching, load adjustments, and changes in control strategies. Process data from different conditions exhibit systematic differences in statistical characteristics, variable coupling relationships, and dynamic response features, resulting in a significant domain shift problem. Deep learning early warning models trained based on the independent and identically distributed assumption often experience performance degradation, such as a sharp drop in prediction accuracy and a significant increase in false negative / false positive rates, when deployed to new conditions not seen during training.
[0047] Third, current multi-condition processing methods all have key limitations. There are currently three main processing paths: condition-specific modeling strategies rely heavily on reliable condition identification logic, which can easily lead to false alarms during condition switching transitions, and cannot be directly applied to new conditions not seen during the training phase; domain adaptation methods usually require access to target domain data for alignment, and this assumption often does not hold true for newly put into operation conditions, extreme conditions, or scenarios with scarce historical data; domain generalization methods only use multiple known condition data to train the model, but related research mainly focuses on fault diagnosis classification tasks, and systematically extending it to fault early warning multi-step sequence prediction tasks is still a blank.
[0048] Fourth, the generalization across operating conditions under fault early warning faces more severe challenges. Compared with the classification task in fault diagnosis, fault early warning requires consistent prediction accuracy and early warning stability across multiple time domains: the evolution trajectory of the same fault under different operating conditions may have significant differences in amplitude, delay time, and variable coupling strength; multi-step prediction is prone to error accumulation under closed-loop control dynamics and operating condition transition disturbances, which rapidly weakens the reliability of early warning; early warning results must be mapped to executable alarm decision criteria, making assessment and deployment more sensitive to the dynamic characteristics dependent on operating conditions.
[0049] In summary, how to achieve domain-invariant learning of multi-step prediction models for chemical processes under multiple operating conditions, so that they can be directly deployed without any retraining in completely unseen operating conditions, is a core technical problem that urgently needs to be solved in this field.
[0050] To address the aforementioned technical issues, this disclosure provides a cross-condition fault early warning method for chemical processes based on Hierarchical Domain-Invariant Alignment Networks (HIAN). This method models the multi-condition fault early warning task as a multi-source domain generalization problem. By hierarchically aligning the feature distributions of multiple known source conditions, it drives the HIAN, i.e., the prediction model, to learn a dynamic representation of the process that is insensitive to changes in operating conditions. This enables direct deployment of zero samples for completely unseen operating conditions without requiring any target operating condition data or retraining or fine-tuning for new operating conditions. Consequently, it solves the problem of decreased generalization performance of prediction models in multi-condition chemical processes due to changes in operating conditions (domain shift). Furthermore, this method employs a newly designed backbone prediction network based on a Transformer encoder-decoder architecture. A residual domain adaptation module is added after each encoder layer to convert sequence-level hidden state into sample-level aligned features. An adaptive layer-wise weighting (ALW) mechanism is introduced to automatically assign alignment strength to each encoder layer. A hierarchical multi-kernel maximum mean difference (MK-MMD) alignment loss is used to simultaneously measure and minimize the feature distribution differences between different source conditions across multiple encoder layers. Further, in the offline modeling stage, a dual-stream training strategy is employed to jointly optimize the prediction loss and hierarchical alignment loss. In the online early warning stage, the trained model parameters are loaded, and zero-sample rolling multi-step prediction with only forward propagation is performed on the real-time data of the target condition. The predicted trajectory is compared with the alarm limit to achieve fault early warning. This method can be deployed directly without any target operating condition data, significantly reducing the modeling cost of the early warning system for new operating conditions. Hierarchical feature alignment effectively captures multi-scale domain shifts, and the residual domain adaptation module balances transferability and prediction accuracy. It has been verified on Tennessee-Eastman process benchmarks and industrial fluidized catalytic cracking unit data that it has significantly better cross-operating condition generalization ability and early warning performance than existing methods.
[0051] Unlike fault diagnosis and classification tasks (which output fault category labels), the multi-step sequence prediction task in this embodiment is a multivariate time-series regression task, outputting continuous values of key process variables at multiple future time nodes / moments. Furthermore, the early warning decision in this embodiment compares these continuous values with predefined alarm conditions, triggering an early warning when the predicted value exceeds a limit.
[0052] Now combined Figures 1 to 8This disclosure provides an illustrative description of a cross-condition fault early warning method for chemical processes. This method may include two parts: an offline training / modeling method and an online zero-shot early warning method. For example... Figure 1 As shown, the online zero-sample early warning method may include the following steps S101 to S103.
[0053] In step S101, the first process data stream of the target process flow under the target operating condition is obtained.
[0054] The target operating condition refers to a completely unseen operating condition, for which related data has not been used in the training of the predictive model. Specifically, the target operating condition refers to a new operating condition that is not involved in the offline training phase and whose data is not visible throughout the entire modeling process; this operating condition corresponds to a zero-sample deployment scenario. The trained predictive model provided in this embodiment can output multi-step prediction results for the target operating condition.
[0055] The target process flow refers to the process flow that is of concern in actual chemical processes, and the specific flow depends on the actual situation.
[0056] The first process data stream can include observations of each first process variable at multiple consecutive first historical moments. This first process data stream can be acquired in real time through a distributed control system (DCS) actually deployed in the chemical process environment. The sampling interval is consistent with the second process data stream used in the offline modeling phase. Furthermore, the dimensions, order, and units of the process variables in both the offline training and online zero-sample early warning phases must be strictly aligned. First process variables refer to key process variables in the chemical process. Key process variables are core measurable physical / chemical parameters that are prioritized for real-time monitoring and early warning, such as reactor temperature, pressure, flow rate, and composition. Since there are multiple first process variables, the first process data stream includes observations of multiple first process variables at multiple consecutive first historical moments.
[0057] In step S102, the preprocessed first process data stream is input into the trained prediction model to obtain the prediction result.
[0058] The prediction results can include predicted values for each first process variable at multiple future times. Since there are multiple first process variables, the prediction results can include predicted values for multiple first process variables at multiple future times. For example, if the input to the prediction model is real-time process data for the two first process variables, pressure and reactor temperature, then the prediction model will output prediction results for pressure and prediction results for reactor temperature. The pressure prediction result will include predicted pressure values at multiple future times, and the reactor temperature prediction result will include predicted temperature values at multiple future times. Figure 2aAs shown, predictions are made based on the observed values of the first process variable 1, the first process variable 2, ..., the first process variable N at L historical times, and the predicted values of these process variables at H future times are obtained.
[0059] In short, the prediction model provided by this method adopts a multi-input multi-output (M-in / M-out) architecture. Multiple inputs refer to the historical observations of multiple process variables, while multiple outputs are the predicted values of these process variables at future times. This method employs a joint supervision design of M-in / M-out, and synchronous supervision of multiple variables helps the encoder learn the joint distribution of process dynamics independent of operating conditions, improving generalization. Furthermore, the same model can provide future predictions for multiple candidate key variables, offering flexibility in deployment.
[0060] Preprocessing of the first process data stream may include: obtaining the standardized parameters corresponding to each first process variable. The standardized parameters are the mean and standard deviation calculated based on the observations of a second process variable (the same as the first process variable) at multiple second historical moments. The observations of the second process variable at multiple second historical moments constitute the second process data stream used for training the prediction model, as described later. The process for determining the standardized parameters is explained below. Standardizing each observation of the first process variable according to the standardized parameters corresponding to each first process variable is also described. For example, for the first process variable of pressure, its first process data streams p1, p2…pn are obtained, and the data streams for… The average pressure μ1 and standard deviation σ1 (μ1 and σ1 are calculated using pressure observations at multiple second historical moments) are given. For observation p1, the standardized observation p1' can be obtained by p1' = (p1 - μ1) / σ1. The standardization of p2...pn is similar to the standardization of p1'. The standardization of other first process variables is similar to that of pressure. For each first process variable, the standardized observations at a first number of first historical moments are used as the input matrix. For example, when the real-time acquired first process data stream arrives at time t, the standardized observations from the most recent L times (t-L+1) to t can be used to form the input matrix. The input matrix can be denoted as... , This represents the standardized observation value at time t-L+1, the first historical moment. Let t represent the standardized observations at the first historical moment, t; M represents the number of first process variables; the input matrix serves as the input to the trained prediction model; and the prediction model evaluates the input... After processing, the predicted values of these M first process variables at multiple future times can be obtained. In this way, the preprocessing of the first process data stream reuses standardized parameters from historical operating condition data, achieving rapid alignment of cross-operating condition data distribution. The fixed-length input matrix design preserves the temporal correlation of process variables and unifies the model input dimension, enabling the trained model to directly adapt to the target operating condition. This preprocessing method shortens the model deployment cycle while ensuring data consistency, and improves the real-time performance and robustness of cross-operating condition fault early warning.
[0061] In step S103, an early warning scheme for the target working condition is determined based on the prediction results and preset early warning conditions.
[0062] Each primary process variable has a corresponding process safety limit. These limits can be sequentially lowered to include the high-high alarm limit (HH), high alarm limit (H), low alarm limit (L), and low-low alarm limit (LL). The high-high and low-low alarm limits are hazard level parameters. When the value of a primary process variable exceeds the high-high alarm limit or falls below the low-low alarm limit, the chemical process's operating data has reached a critical hazard level, indicating an impending safety accident. In this case, an alarm / warning must be issued, referring to [reference needed]. Figure 2b The high and low reporting limits are warning level parameters. When the value of the first process variable exceeds the high reporting limit but does not reach the high-high reporting limit, or falls below the low reporting limit but does not fall below the low-low reporting limit, the operating data of the chemical process exceeds the normal operating range but has not yet reached the critical danger level. The specific values of the process safety limits can be flexibly set according to the actual situation.
[0063] The early warning conditions include those for a target process variable among multiple first process variables. The target process variable refers to a pre-specified key monitoring variable, i.e., one of the multiple first process variables. The system determines whether the predicted value of the target process variable at a future time meets the early warning conditions for that target process variable. If so, a fault alarm is triggered / activated.
[0064] Step S103 may include: if the predicted value of a target process variable at a future time is less than a preset first threshold, or if the predicted value of a target process variable at a future time is greater than a preset second threshold, then the early warning scheme for the target operating condition includes activating a fault alarm for the target process variable. The first threshold is lower than the second threshold.
[0065] The early warning process in this embodiment can have two scenarios. In the first scenario, the first threshold can be a low-low alarm limit and the second threshold can be a high-high alarm limit. In this case, if the predicted value of the target process variable at a future time is less than the low-low alarm limit or the predicted value of the target process variable at a future time is greater than the high-high alarm limit, it indicates that the operating data of the chemical process has reached a critical danger level, and the early warning scheme can be determined to activate a fault alarm for the target process variable. Otherwise, the operating data of the chemical process can be considered to be within the normal range, and no fault alarm will be activated.
[0066] In the second scenario, the first threshold can be the low alarm limit and the second threshold can be the high alarm limit. In this case, if the predicted value of the target process variable at a future time is less than the low alarm limit or the predicted value of the target process variable at a future time is greater than the high alarm limit, it indicates that the operating data of the chemical process is outside the normal operating range. The early warning scheme can be determined to activate the fault alarm for the target process variable. Otherwise, the operating data of the chemical process can be considered to be within the normal range, and no fault alarm will be triggered.
[0067] Thus, the early warning scheme provided in this embodiment monitors the range of predicted values of target process variables at future times by setting dual thresholds. Once the predicted value of the target process variable exceeds the threshold range, a fault alarm is immediately activated for that variable, achieving accurate location of the fault variable and avoiding false alarms or missed alarms. This targeted alarm mechanism can quickly pinpoint the source of the anomaly, reduce fault investigation time, and improve the efficiency and accuracy of safety monitoring of chemical processes.
[0068] like Figure 3 As shown, the offline modeling method (i.e. the training process of the prediction model) may include the following steps S201 to S202.
[0069] In step S201, a training sample set is constructed based on the second process data stream of the sample process flow under multiple sample operating conditions.
[0070] Each sample operating condition differs from the target operating condition. A sample operating condition refers to a known source operating condition, also known as a training condition. The target operating condition and the training operating condition belong to the same chemical process, differing only in operating conditions. The sample process flow refers to the process flow of interest in the actual chemical process; the specific flow depends on the actual situation. In this embodiment, the sample process flow can be the same as the target process flow. Therefore, multi-step prediction of process variables for the same process flow will be more accurate.
[0071] The second process data stream comprises observations of each second process variable at multiple second historical moments. This second process data stream can be acquired through a distributed control system (DCS) actually deployed in the chemical process environment, with sampling intervals consistent with the first process data stream used in the online prediction phase. Second process variables also refer to key process variables in the chemical process. The meaning of key process variables has been explained above and will not be repeated here. For example, a second process variable could be reactor temperature, pressure, flow rate, etc. There can be multiple second process variables. These second process variables may be the same as or different from the first process variables, but preferably, multiple second process variables include multiple first process variables.
[0072] This method offers the advantage of zero-sample cross-operating-condition deployment and requires no target operating-condition data. It constructs multi-operating-condition fault early warning as a multi-source domain generalization problem. The prediction model is trained only on known source operating-condition data and can be directly deployed to completely new operating conditions without any target operating-condition data. For example… Figure 4 The time-series patterns of different operating conditions and the switching relationship between operating conditions are shown. The prediction model learns the evolution of the same process variable under three known source operating conditions (Mode1, Mode2, and Mode3) to predict the change of the process variable under the unknown operating condition (Mode4). This significantly reduces the data collection and modeling costs of the early warning system for new operating conditions.
[0073] Step S201 may include: calculating the mean and standard deviation of the second process variable based on all observed values corresponding to each second process variable. Taking process variable a as an example, its n observed values a1, a2...an under all sample working conditions can be used to calculate the mean μ2 of process variable a through μ2=(a1+a2+...+an) / n, and can be used through σ2= Calculate the standard deviation σ² of process variable a. Based on the mean and standard deviation of each second process variable, standardize each observation of the second process variable. Taking the observation a² of process variable a as an example, the standardized observation a²' can be obtained by a²' = (a² - μ²) / σ². The standardization of observations of other process variables is similar to a², and will not be elaborated further. For each second process variable, use the standardized observations at the target time and the first number of second historical times before the target time as historical samples, and the standardized observations at the second number of second historical times after the target time as prediction samples to construct a training sample set. For example, a sliding window method can be used to construct historical-prediction sample pairs, i.e., sample pairs. The training sample set includes multiple sample pairs. Each sample pair includes historical samples and prediction samples for the same target time, with the number of target times preset to multiple.
[0074] For example, multivariate time-series process data from K known source conditions can be collected. Standardized parameters are estimated and standardized using only the source domain training data. A sliding window approach, or by selecting multiple different target times, is used to construct history-prediction sample pairs to build a training sample set. Multivariate time-series process data refers to the values of multiple second process variables at multiple second historical times. The dataset for the k-th source condition is S^(k)={(X_i^(k), y_i^(k))}_{i=1}^{N_k}, where x_t∈R^M represents the observation vector of M second process variables at time t, which is T×M in matrix form, where T is the number of sampling time points and M is the numerical matrix of the second process variables. An anchor point t (i.e., the target time) is slid across the multivariate time series data of each source condition with a step size of 1. Each anchor point corresponds to a sample pair. Historical samples can be represented as X_t=[x_{t-L+1},…,x_t]∈R^{L×M}, which is the M-dimensional multivariate observation sequence of L time points (i.e., the first quantity) before and after anchor point t, where L is the length of the historical window. Predicted samples can be represented as Y_t=[x_{t+1},…,x_{t+H}]∈R^{H×M}, where H is the prediction step size, i.e., the second quantity. That is, the observations of L+H consecutive time points within the interval [t-L+1,t+H], with the first L as historical samples and the last H as predicted samples, i.e., prediction labels. Furthermore, to avoid introducing heterogeneous dynamics through cross-condition boundary samples, a "condition boundary safety" constraint is adopted, only accepting anchor points t that satisfy m_{t-L+1}=m_{t-L+2}=…=m_{t+H]. This means that the interval covered by a sample pair falls within the same steady-state condition segment; in other words, constructing a sample pair is for data from a single condition and cannot include data from multiple conditions. Additionally, multivariate time-series process data can be divided into training, validation, and test sets in the source domain. The validation and test sets, as well as the first-process data stream, do not participate in the estimation of standardized parameters. Standardization is uniformly applied to the source domain training / validation / test sets and the first-process data stream in the online early warning stage to achieve consistent dimensions across conditions.
[0075] In this way, by calculating the mean and standard deviation of variables through historical data across operating conditions, standardization is used to eliminate differences in dimensions. Then, sample pairs are constructed according to two time windows: historical time window and future time window. This ensures the consistency of data distribution and provides the model with prediction labels / supervision signals for multi-operating condition time series prediction tasks. This makes the training sample set both generalizable and prediction-oriented, laying the data foundation for the model's cross-operating condition fault early warning capability.
[0076] In step S202, the initial prediction model is trained based on the training sample set to obtain the trained prediction model.
[0077] During the offline modeling phase, the hierarchical domain-invariant alignment network (HCI) is configured, which forms the network structure of the prediction model. The prediction model includes a backbone prediction network based on an encoder and decoder, and residual domain adaptation (RDA) modules following each encoder / self-attention layer. The prediction model employs an adaptive layer-wise weighting (ALW) mechanism. During the offline modeling phase, the model's structural parameters are set, and the connection parameters are randomly initialized.
[0078] The encoder may include multiple self-attention layers (or encoder layers) connected in sequence. The self-attention layers are used to extract features from the input to obtain hierarchical features. The input to the self-attention layer includes the input embedding sequence or the hierarchical features output from the previous self-attention layer. The hierarchical features represent the sequence-level hidden state. The encoder extracts hierarchical features progressively through N stacked layers, generating layer-by-layer representations, advancing layer by layer until the Nth layer, obtaining the final encoder output, i.e., the hierarchical features output from the Nth self-attention layer, for use by the decoder for cross-attention. Exemplarily, each encoder layer in this embodiment is a standard Transformer encoder sublayer structure, which may include multi-head self-attention (MSA), a feedforward network (FFN), residual connections, and layer normalization (LN). N such encoder layers are stacked in sequence to form the encoder in the backbone prediction network.
[0079] Each encoder layer is connected to a residual domain adaptation module. The residual domain adaptation module can be used to transform the hierarchical features output by the encoder layer it is connected to into sample-level aligned features. For example, for N encoders, the sequence-level hidden state H^(l)∈R^{L×B×d} (where L represents the time step, B represents the batch size, and d represents the model dimension) output by the first encoder layer is transformed by the corresponding RDA module into sample-level aligned features Z^(l)∈R^{B×d_h} for that layer. Each batch of samples yields a d_h-dimensional feature vector. The transformation process can specifically include three steps: temporal mean pooling, two-layer multilayer perceptron (MLP) transformation, linear residual shortcut addition, and activation. Temporal mean pooling can be expressed as... The two-layer MLP transform (including LN and Dropout) can be expressed as: W_1^(l)∈R^{d×d_h}, W_2^(l)∈R^{d_h×d_h}; the linear residual shortcut addition and activation can be expressed as Among them, residual shortcut The purpose is to retain discriminative information useful for process variable prediction while learning domain-invariant transformation. N encoder layers correspond to N sample-level alignment features {Z^(l)}_{l=1}^N, which are used to align the multi-kernel maximum mean difference (MK-MMD) distribution across source conditions at each layer. This method balances transferability and prediction accuracy through the residual domain adaptation module. The residual connections in the RDA module ensure that the discriminative information required for the prediction task is not discarded while performing domain-invariant transformation, avoiding the dilemma of alignment gain versus prediction loss in traditional domain adaptation methods.
[0080] For example, the encoder may include a first self-attention layer, an intermediate self-attention layer, and a final self-attention layer connected in sequence. The first self-attention layer is followed by a first residual domain adaptation module, the intermediate self-attention layer by a second residual domain adaptation module, and the final self-attention layer by a third residual domain adaptation module. The input of the first self-attention layer is the input embedding sequence, and the output is a first hierarchical feature. The input of the first residual domain adaptation module is the first hierarchical feature, and the output is a first sample-level alignment feature. The input of the intermediate self-attention layer is the first hierarchical feature, and the output is a second hierarchical feature. The input of the second residual domain adaptation module is the second hierarchical feature, and the output is a second sample-level alignment feature. The input of the final self-attention layer is the second hierarchical feature, and the output is a third hierarchical feature. The input of the third residual domain adaptation module is the third hierarchical feature, and the output is a third sample-level alignment feature.
[0081] In this way, by extracting hierarchical features through a self-attention layer and combining it with a residual domain adaptation module to achieve sample-level feature alignment, the problem of data distribution differences across operating conditions is effectively solved. The encoder structure enhances feature representation capabilities, improves the model's fault prediction accuracy for complex chemical processes, provides more reliable feature support for fault early warning across operating conditions, and reduces false alarms and missed alarms caused by data distribution offsets.
[0082] This method may further include: performing a linear transformation on the historical input matrix corresponding to all historical samples in the training sample set to obtain a first matrix, where the historical input matrix / sequence refers to the historical samples... It contains standardized numerical matrices of M different second process variables at L consecutive second historical moments; the first matrix is positionally encoded to obtain the input embedding sequence.
[0083] Linear transformation / projection refers to mapping the vector representation of each element in an input sequence from an original feature space to a new feature space through a learnable linear transformation. For example, a historical input sequence... Each row, i.e., the M-dimensional multivariate observation vector at each second historical moment or historical time point, represents the M-dimensional multivariate observation vector. (Not a single scalar, nor a single variable), the specific form of linear projection is: Where W_emb∈R^{M×d} and b_emb∈R^d are learnable parameters, and each time point will... Mapping from the original variable dimension M to the model dimension d, we obtain the t-th row of H^(0), followed by layer normalization (LN) to stabilize training across operating conditions.
[0084] Position encoding can be achieved using the standard sinusoidal position encoding formula defined by Vaswani et al. (2017): PE_{(pos,2i)}= sin(pos / 10000^{2i / d}), PE_{(pos,2i+1)} = cos(pos / 10000^{2i / d}), where pos represents the position index of the data in the sequence (such as the 1st or Lth time step), i is the dimension index, and d is the dimension of the feature vector. For the 2i-th dimension (even dimension) of position pos, a sine function is used to generate the encoding, and for the 2i+1-th dimension (odd dimension), a cosine function is used. Both share the same frequency parameter 10000^{2i / d}. This design ensures that the encoding of different positions is unique. Low-frequency components (such as low-dimensional i) capture the global temporal structure, while high-frequency components (such as high-dimensional i) reflect local details. The position index pos=1,…,L is encoded into a d-dimensional vector PE_pos, and then superimposed on the corresponding row of H^(0) according to the position (H^(0)_t ← H^(0)_t + PE_t), thereby injecting temporal position information into the self-attention layer, enabling the self-attention layer to perceive the temporal order of chemical process data.
[0085] In this way, the historical input matrix is mapped to a high-dimensional space through linear transformation, which enhances the feature expression capability. Combined with positional encoding, the temporal information of the time series is preserved, providing ordered input for the self-attention layer and laying the foundation for subsequent hierarchical feature extraction. This enables the prediction model to more accurately capture the correlation and dynamic changes between variables, and improve the accuracy and robustness of cross-operating condition fault early warning.
[0086] The prediction model trained and initialized based on the training sample set in step S202 may include: calculating the hierarchical multi-kernel maximum mean difference value based on the sample-level alignment features output by all residual domain adaptation modules. The multi-kernel maximum mean difference is a statistic used to measure the difference between two probability distributions, widely applied in transfer learning to align the feature distributions of the source domain (known source operating condition data) and the target domain (target operating condition data). For example, referring to... Figure 5 For the two source cases si and sj, based on the sample-level alignment feature Z output by the RDA module following the Encoder Layer 1 for si... (1)The sample-level alignment feature Z of the RDA module output after the Encoder Layer 1 for sj (1) Calculate the hierarchical multi-core maximum mean difference (MMD Loss) for the corresponding Encoder Layer 1 for si and sj. If there are n Encoder Layers, then there are n maximum mean difference values for the multi-kernel layer for si and sj. The maximum mean difference values for the multi-kernel layer of the remaining layers are similar to those of Encoder Layer 1, and will not be elaborated further. Furthermore, the mean squared error (MSE) value is calculated based on the sample-level alignment features output by all residual domain adaptation modules and the predicted samples in the training sample set. For example, the Decoder for si obtains the prediction result Prediction i based on all sample-level alignment features and related attention information, and the Decoder for sj obtains the prediction result Prediction j based on all sample-level alignment features and related attention information. The MSE Loss is calculated based on Prediction i and the corresponding predicted sample, and Prediction j and the corresponding predicted sample. For example, Prediction i in the corresponding predicted sample can be the predicted value of the second process variable a at the H times t+1,…,t+H, obtained based on the standardized observations of the L second historical times t-L+1,…,t for the second process variable a. i and the corresponding predicted sample in the corresponding predicted sample can be the standardized observations of the second process variable a at the H second historical moments t+1,…,t+H; the hierarchical multi-kernel maximum mean difference (MK-MMD) and mean squared error are used as the objective function value, and the parameters of the encoder are adjusted by minimizing the objective function value. Figure 5 As shown, this method employs a dual-stream training strategy, simultaneously training from two different source conditions (e.g., ...). Figure 5 In the source conditions (Source Mode si and Source Mode sj), small batches of data are sampled and forward-propagated through the shared backbone prediction network. The mean squared error (prediction loss) is then calculated. Figure 5 MSE Loss in hierarchical multi-kernel maximum mean difference (alignment loss, such as...) Figure 5 MMD Loss MMDLoss ... MMD Loss The overall objective function composed of (e.g.) Figure 5The method employs a Total Loss mechanism to jointly optimize encoder parameters. This training method constructs a loss function by fusing hierarchical multi-kernel maximum mean difference (MK-MMD) and mean squared error (MSE), enabling the model to extract general features across operating conditions while maintaining prediction accuracy. Hierarchical MK-MMD aligns feature distributions across different operating conditions, reducing inter-domain differences; MSE constrains the error between predicted and true values, improving time-series prediction capabilities. The synergistic optimization of these two mechanisms makes the model more stable and accurate in cross-operating-condition scenarios of chemical processes, providing a reliable feature foundation for fault early warning. This method, through hierarchical MK-MMD alignment, enables the prediction model to learn process dynamic representations that are insensitive to changes in operating conditions, thus maintaining prediction accuracy even for unseen target operating conditions.
[0087] The prediction model employs an adaptive layer-wise weighting (ALW) mechanism for parameter tuning. Adjusting the encoder parameters includes updating the alignment weights of each self-attention layer based on the sample-level alignment features output by the residual domain adaptation module. These alignment weights are used to adjust the alignment strength of the self-attention layers. During model training, the learnable parameter vector is jointly optimized with the backbone prediction network parameters through backpropagation. The weight allocation for each layer is automatically learned based on the hierarchical alignment loss (i.e., the maximum mean difference between hierarchical multi-kernel values) calculated from the sample-level alignment features output by the residual domain adaptation module. This data-driven, automatically adjusted weight allocation method allows self-attention layers with high transferability to receive greater weights, while shallower, condition-specific layers receive smaller weights. The learnable parameter vector is determined and fixed after the offline training phase. During the online alert / inference phase, this parameter vector is directly loaded and not updated further; it remains the same for all samples during inference. For example, a learnable parameter vector w = [w_1,…,w_N] ∈ R^N is introduced, where w_1 represents the alignment weight of the first self-attention layer, w_N represents the alignment weight of the Nth self-attention layer, and N is the number of self-attention layers. The alignment weights of each layer are generated through a normalized exponential function (softmax) and smoothing, automatically adjusting the alignment strength of each layer. Specifically, each component of the learnable parameter vector w is first fed into the function α_l = exp(w_l) / Σ_{k=1}^Nexp(w_k), and then smoothed by α_l ← (1-δ)·α_l + δ / N, where δ ∈ (0,1) is a small smoothing coefficient to prevent weight collapse (ensuring that the weights of each layer do not approach zero and become "inactive"). Finally, {α_l} satisfies α_l > 0 and Σ_{l=1}^Nα_l = 1, which represents the alignment weights of the N self-attention layers. This method avoids negative migration through an adaptive layer-by-layer weighting mechanism. By jointly optimizing learnable weights, it automatically identifies and strengthens the alignment of transferable representation layers while suppressing over-alignment of condition-specific layers, effectively avoiding the risk of negative migration.
[0088] This method deploys an RDA module after each layer of the encoder and combines it with the ALW mechanism to adaptively weight the multi-layer alignment loss. It can perform comprehensive, multi-scale feature alignment from low-level statistical offsets to high-level dynamic mode differences, effectively capturing multi-scale domain offsets. Ablation experiments verify that hierarchical alignment reduces the MSE by more than two orders of magnitude compared to single-layer alignment in unseen conditions.
[0089] During the training of the prediction model, the decoder generates a predicted sequence for the next H steps based on the encoder output. It then uses the AdamW optimizer combined with cosine annealing learning rate scheduling for optimization, maintaining the exponential moving average (EMA) model parameters. The encoder output refers to the hierarchical features of the last layer (i.e., the Nth self-attention layer) among multiple self-attention layers, i.e., the sequence-level hidden state H^(N)∈R^{L×d}. The decoder generates the predicted sequence for the next H steps through cross-attention, using H^(N) as the key / value memory and combining it with its own masked self-attention. (or key variable form) This refers to the predicted values of process variables at multiple future times. The source domain validation set data is input into the trained prediction model, and the mean squared error of multi-step predictions is calculated to evaluate the model's fitting ability and generalization ability across operating conditions.
[0090] Thus, after completing the training process of the prediction model, the trained prediction model is obtained. During the online early warning process, real-time data stream acquisition and standardization preprocessing are performed first. The first process data stream of the target operating condition is acquired in real time from the distributed control system, and the standardized parameters estimated in the offline stage are used to standardize the first process data stream. Then, rolling multi-step forward prediction is performed. A historical input window is constructed using a sliding window method, the EMA model parameters obtained from offline training are loaded, and forward propagation inference is executed only. A rolling autoregressive inference strategy is used to generate multi-step prediction trajectories of the target operating condition, i.e., the predicted values of the first process variables at multiple future times. Finally, fault early warning decision-making and alarm output are executed. The multi-step prediction trajectory is compared with predefined process safety limits. If the predicted value exceeds the alarm limit, a fault early warning signal is triggered. Alarm events are merged and performance is evaluated according to alarm management standards.
[0091] Based on Figure 6Taking the multi-condition zero-sample fault prediction of the Tennessee Eastman Process (TEP) benchmark as an example, this embodiment addresses the fault prediction of a multi-condition chemical process. The main problem it solves is the decline in the generalization performance of the prediction model due to domain shift in data collected under different conditions. The dataset used in this embodiment adopts the TEP simulation benchmark revised by Bathelt et al., which covers seven steady-state conditions (conditions 0-6). Each condition exhibits systematic differences in product G / H molar ratio, reactor level / temperature / pressure, and valve configuration. This embodiment uses fault 13 (slow reaction kinetic drift) as the experimental scenario, selecting five variables from the revised TEP process variables that are most closely physically coupled with the reactor-separator subsystem to construct the prediction model. The data sampling interval is one minute.
[0092] Experimental setup: Conditions 0, 1, and 2 are the source domain (training / validation); conditions 3, 4, 5, and 6 are unseen target domains (for testing only). Each condition package has 10 independent simulation batches: batches 1-8 are used for training, batch 9 for validation, and batch 10 for testing. Prediction model configuration: 3-layer Transformer encoder, hidden layer dimension d=64, 4 attention heads, MK-MMD alignment weight λ=5.0, history window L=32, prediction stride H=32. MK-MMD uses 5 RBF kernels with different bandwidths, and EMA momentum coefficient ρ=0.999. Optimizer: AdamW, initial learning rate 10. -4 Cosine annealing scheduling. Experimental results: In zero-sample generalization evaluation with conditions 0, 1, and 2 as the source domain and conditions 3, 4, 5, and 6 as the completely unseen target domain, HIAN reduces the mean fault prediction MSE by approximately 75% compared to LSTM and by approximately 80% compared to Vanilla Transformer on unseen conditions, significantly narrowing the generalization gap. Ablation experiments further confirm that hierarchical alignment is a core component of HIAN: after completely removing alignment constraints, the MSE of unseen conditions increases by more than two orders of magnitude.
[0093] This example uses zero-sample cross-condition fault early warning based on industrial FCC unit data. The dataset used in this embodiment is collected from 170 days of continuous operation data (1-minute sampling, approximately 244,800 time points) of a fluidized catalytic cracking (FCC) unit in a petrochemical enterprise. The target variable for early warning is the regenerator flue gas outlet temperature. Input variables include four variables: dense bed temperature, fresh feed flow rate, riser reaction temperature, and flue gas O2 content. The model outputs predicted values for these five variables for the next H steps. The predicted column for the pre-specified key monitoring variable, regenerator flue gas outlet temperature, is compared with predefined high-high / low-low alarm thresholds to trigger an early warning. The predicted columns for the other four variables are not involved in the alarm decision in this embodiment. Through unsupervised clustering of daily operating characteristics (UMAP dimensionality reduction + Gaussian mixture model GMM), eight steady-state operating conditions are identified, such as... Figure 7 As shown. Experimental setup: Operating conditions 0, 1, and 4 were used as the source domain, and operating conditions 2, 3, 5, 6, and 7 were used as the unseen target domain. The model configuration was exactly the same as the TEP experiment. Alarm performance evaluation adopted a 5-step advance prediction (5-minute warning window), and the alarm threshold was set to ±2σ.
[0094] according to Figure 8 The comparison of the 5-step advance prediction results for FCC unseen condition 6 shows that, compared to Transformer and LSTM, the HIAN provided in this embodiment has no false alarms (FP). FP refers to false positives, which incorrectly predict the number of negative classes as positive. TP refers to true positives, which correctly predict the number of positive classes as positive. FN refers to false negatives, which incorrectly predict the number of positive classes as negative. Table 1 shows a comparison of the prediction accuracy of different models (average MSE for unseen condition). As can be seen from Table 1, the HIAN provided in this embodiment reduces the MSE of the unseen condition by 12.7% compared to LSTM and by 18.7% compared to Transformer. Ablation analysis shows that the independent contribution of the dual-stream training structure is small (0.25% improvement), and explicit MK-MMD hierarchical alignment is the main source of improvement (3.21% improvement). Table 2 shows a comparison of the alarm performance of different models (unseen condition, alarm threshold α=2.0). As can be seen from Table 2, the HIAN provided in this disclosure achieves the highest early warning accuracy (PA=85.8%) with the fewest false alarms (FP=41, which is 48% and 49% lower than LSTM and Transformer, respectively). The PA remains above 78% in all 5 unseen operating conditions, proving that the cross-operating condition generalization capability can be directly translated into more reliable industrial early warning performance.
[0095] Table 1. Prediction accuracy of different models
[0096] Table 2 Alarm performance of different models
[0097] As can be seen, the proposed method significantly outperforms the baseline method on the multi-condition TEP benchmark, and its practical engineering value is verified on industrial FCC device data, proving that the cross-condition generalization capability can be directly translated into more reliable industrial early warning performance.
[0098] In summary, the proposed cross-condition fault early warning method for chemical processes, based on a hierarchical domain-invariant alignment network, achieves zero-sample cross-condition fault early warning for completely unseen conditions. It utilizes a hierarchical Transformer encoder to extract features from multi-source operating condition data, transforms sequence-level representations into sample-level aligned features through a residual domain adaptation module, and combines an adaptive layer-by-layer weighting mechanism with hierarchical MK-MMD alignment loss to achieve zero-sample fault early warning for completely unseen operating conditions. This method demonstrates significantly superior cross-condition generalization ability and early warning performance compared to existing methods on TEP benchmarks and industrial FCC unit data, possessing significant practical implications for improving the safety monitoring level of multi-condition chemical processes.
[0099] This disclosure also provides a cross-condition fault early warning device for chemical processes. The device includes: a data acquisition module for acquiring a first process data stream of a target process under a target condition, the first process data stream including observations of multiple first process variables at multiple first historical moments; a fault prediction module for inputting the preprocessed first process data stream into a trained prediction model to obtain a prediction result, the prediction result including predicted values of the multiple first process variables at multiple future moments; and an early warning generation module for determining an early warning scheme for the target condition based on the prediction result and preset early warning conditions. The device further includes a model training module for executing the training process of the prediction model, the training process including: constructing a training sample set based on second process data streams of sample process flows under multiple sample conditions, the second process data stream including observations of each second process variable at multiple second historical moments, each sample condition being different from the target condition; and training an initialized prediction model based on the training sample set to obtain a trained prediction model.
[0100] In one possible implementation, constructing a training sample set based on the second process data stream of sample process flows under multiple sample operating conditions includes: calculating the mean and standard deviation of the second process variable based on all observations corresponding to each second process variable; standardizing each observation of the second process variable based on the mean and standard deviation of each second process variable; for each second process variable, using the standardized observations at a first number of second historical times before the target time as historical samples, and using the standardized observations at a second number of second historical times after the target time as prediction samples, to construct the training sample set, which includes multiple sample pairs, each sample pair including historical samples and prediction samples for the same target time, wherein the number of target times is preset to be multiple.
[0101] In one possible implementation, the prediction model includes an encoder comprising a plurality of self-attention layers connected in sequence, each self-attention layer being connected to a residual domain adaptation module; the self-attention layers are used to extract features from the input to obtain hierarchical features, the input of the self-attention layer including the input embedding sequence or the hierarchical features output by the previous self-attention layer; the residual domain adaptation module is used to transform the hierarchical features output by the self-attention layers to which it is connected into sample-level aligned features.
[0102] In one possible implementation, the apparatus further includes a sample processing module, configured to: perform a linear transformation on the historical input matrix corresponding to all historical samples of the training sample set to obtain a first matrix; and perform position encoding on the first matrix to obtain the input embedding sequence.
[0103] In one possible implementation, training the initialized prediction model based on the training sample set includes: calculating the hierarchical multi-kernel maximum mean difference value based on the sample-level alignment features output by all residual domain adaptation modules; and calculating the mean squared error value based on the sample-level alignment features output by all residual domain adaptation modules and the prediction samples in the training sample set; using the hierarchical multi-kernel maximum mean difference value and the mean squared error value as objective function values, and minimizing the objective function value as the training objective, adjusting the parameters of the encoder.
[0104] In one possible implementation, adjusting the parameters of the encoder includes: updating the alignment weights of each self-attention layer based on the sample-level alignment features output by the residual domain adaptation module, wherein the alignment weights are used to adjust the alignment intensity of the self-attention layers.
[0105] In one possible implementation, the preprocessing of the first process data stream includes: obtaining standardized parameters corresponding to each first process variable, wherein the standardized parameters refer to the mean and standard deviation calculated based on the observations of the same second process variable at multiple second historical moments; standardizing each observation of the first process variable according to the standardized parameters corresponding to each first process variable; and for each first process variable, using the standardized observations at a first number of first historical moments as an input matrix, wherein the input matrix is used as the input of the trained prediction model.
[0106] In one possible implementation, the warning conditions include warning conditions for a target process variable among the plurality of first process variables; determining the warning scheme for the target operating condition based on the prediction result and the preset warning conditions includes: if the predicted value of the target process variable at the future time is less than a preset first threshold, or the predicted value of the target process variable at the future time is greater than a preset second threshold, then the warning scheme for the target operating condition includes activating a fault alarm for the target process variable, wherein the first threshold is lower than the second threshold.
[0107] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0108] This disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above method.
[0109] This disclosure also provides a non-volatile computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the above-described method.
[0110] This disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method.
[0111] Figure 9 A block diagram of a cross-condition fault early warning device for chemical processes provided in an embodiment of this disclosure is shown. For example, device 1900 can be provided as a server or terminal device. (Refer to...) Figure 9The apparatus 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by memory 1932 for storing instructions, such as application programs, that can be executed by the processing component 1922. The application programs stored in memory 1932 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 1922 is configured to execute instructions to perform the methods described above.
[0112] Device 1900 may also include a power supply component 1926 configured to perform power management of device 1900, a wired or wireless network interface 1950 configured to connect device 1900 to a network, and an input / output interface 1958 (I / O interface). Device 1900 can operate on an operating system, such as Windows Server, stored in memory 1932. TM macOS X TM Unix TM Linux TM FreeBSD TM Or similar.
[0113] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by a processing component 1922 of the device 1900 to perform the above-described method.
[0114] Computer-readable storage media can be tangible devices capable of holding and storing programs / instructions used by instruction execution devices. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0115] The computer program (or computer-readable program instructions) described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage medium in the respective computing / processing device.
[0116] The computer program (or computer program instructions) used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing state information from the computer-readable program instructions to implement various aspects of this disclosure.
[0117] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0118] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0119] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0120] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0121] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for early warning of faults across operating conditions in a chemical process, characterized in that, The method includes: Acquire the first process data stream of the target process flow under the target operating condition. The first process data stream includes the observation values of multiple first process variables at multiple first historical moments. The preprocessed first process data stream is input into the trained prediction model to obtain the prediction result, which includes the predicted values of the multiple first process variables at multiple future times. Based on the prediction results and the preset warning conditions, a warning scheme is determined for the target working condition; The method further includes a training process for the prediction model, which includes the following steps: A training sample set is constructed based on the second process data stream of the sample process flow under multiple sample working conditions. The second process data stream includes the observation value of each second process variable at multiple second historical moments. Each of the sample working conditions is different from the target working condition. The initial prediction model is trained based on the training sample set, and the trained prediction model is obtained.
2. The method according to claim 1, characterized in that, The construction of the training sample set based on the second process data stream of sample process flow under multiple sample operating conditions includes: The mean and standard deviation of the second process variable are calculated based on all observations corresponding to each second process variable; Each observation of the second process variable is standardized based on the mean and standard deviation of each second process variable; For each of the second process variables, the standardized observations at a first number of second historical moments before the target time are used as historical samples, and the standardized observations at a second number of second historical moments after the target time are used as prediction samples, so as to construct the training sample set. The training sample set includes multiple sample pairs, and each sample pair includes historical samples and prediction samples for the same target time. The number of target times is preset to be multiple.
3. The method according to claim 1, characterized in that, The prediction model includes an encoder, which includes multiple self-attention layers connected in sequence, and each self-attention layer is connected to a residual domain adaptation module. The self-attention layer is used to extract features from the input to obtain hierarchical features. The input of the self-attention layer includes the input embedding sequence or the hierarchical features output by the previous self-attention layer. The residual domain adaptation module is used to convert the hierarchical features output by the self-attention layer to which it is connected into sample-level aligned features.
4. The method according to claim 3, characterized in that, The method further includes: A linear transformation is performed on the historical input matrix corresponding to all historical samples in the training sample set to obtain the first matrix; The first matrix is positionally encoded to obtain the input embedding sequence.
5. The method according to claim 3, characterized in that, The prediction model trained and initialized based on the training sample set includes: The hierarchical multi-kernel maximum mean difference value is calculated based on the sample-level alignment features output by all residual domain adaptation modules, and the mean square error value is calculated based on the sample-level alignment features output by all residual domain adaptation modules and the predicted samples in the training sample set. The maximum mean difference of the multi-core layer and the mean square error are used as the objective function value. The parameters of the encoder are adjusted to minimize the objective function value as the training objective.
6. The method according to claim 5, characterized in that, The adjustment of the encoder parameters includes: Based on the sample-level alignment features output by the residual domain adaptation module, the alignment weights of each self-attention layer are updated, and the alignment weights are used to adjust the alignment strength of the self-attention layer.
7. The method according to any one of claims 1 to 6, characterized in that, Preprocessing of the first process data stream includes: Obtain the standardized parameters corresponding to each of the first process variables. The standardized parameters are the mean and standard deviation calculated based on the observations of the same second process variable as the first process variable at multiple second historical moments. Each observation of the first process variable is standardized according to the standardization parameter corresponding to each first process variable; For each of the first process variables, the standardized observations at the first historical moment of the first number are used as the input matrix, and the input matrix is used as the input of the trained prediction model.
8. The method according to any one of claims 1 to 6, characterized in that, The warning conditions include warning conditions for the target process variable among the plurality of first process variables; The step of determining an early warning scheme for the target operating condition based on the prediction results and preset early warning conditions includes: If the predicted value of the target process variable at the future time is less than a preset first threshold, or the predicted value of the target process variable at the future time is greater than a preset second threshold, then the early warning scheme for the target operating condition includes activating a fault alarm for the target process variable, wherein the first threshold is lower than the second threshold.
9. A cross-condition fault early warning device for chemical processes, characterized in that, The device includes: The data acquisition module is used to acquire the first process data stream of the target process under the target operating conditions. The first process data stream includes the observation values of multiple first process variables at multiple first historical moments. The fault prediction module is used to input the preprocessed first process data stream into the trained prediction model to obtain prediction results, which include the predicted values of the multiple first process variables at multiple future times. The early warning generation module is used to determine an early warning scheme for the target working condition based on the prediction results and preset early warning conditions. The device further includes a model training module for executing the training process of the prediction model, the training process including: A training sample set is constructed based on the second process data stream of the sample process flow under multiple sample working conditions. The second process data stream includes the observation value of each second process variable at multiple second historical moments. Each of the sample working conditions is different from the target working condition. The initial prediction model is trained based on the training sample set, and the trained prediction model is obtained.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 8.