A mechanism-based hierarchical decoupling method for petroleum refining prediction and attribution

By adopting a mechanism-based topological hierarchical decoupling method, the problems of logical interpretability and dynamic operating condition attribution in petroleum refining modeling are solved, achieving high-precision and transparent prediction and diagnosis, reducing computational complexity, and improving the interpretability and security of the model.

CN121812004BActive Publication Date: 2026-06-02NANJING UNIV OF INFORMATION SCI & TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING UNIV OF INFORMATION SCI & TECH
Filing Date
2026-03-06
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing petroleum refining modeling techniques struggle to balance the accuracy of complex nonlinear fitting, logical interpretability consistent with chemical mechanisms, and dynamic operating condition attribution. Traditional methods cannot effectively capture complex synergistic or antagonistic effects, deep learning models lack transparency, some interpretable methods ignore plant topology leading to rampant spurious correlations, and existing attribution methods lack dynamic time-varying benchmarks.

Method used

A mechanism-based topology hierarchical decoupling method is adopted. By introducing dynamic monotonicity verification and mechanism topology information to restrict interaction terms, a dynamic time-varying benchmark is constructed, generating a structured attribution report that conforms to chemical engineering common sense, eliminating spurious correlation interactions, and ensuring model logical consistency and accurate diagnosis.

Benefits of technology

It achieves high-precision prediction and a transparent model that conforms to chemical engineering common sense, reduces computational complexity, improves the interpretability and security of the model, and can quantify the impact of process parameters on performance in real time, supporting accurate diagnosis and optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121812004B_ABST
    Figure CN121812004B_ABST
Patent Text Reader

Abstract

The application discloses a kind of petroleum refining forecasting and attribution method based on mechanism topological layer decoupling, belong to petroleum chemical industry intelligent modeling and optimization control field, first acquisition original petroleum refining data, and the multi-dimensional feature standardization processing is carried out to data, and the feature data after standardization is obtained;Then based on multimodal competition mechanism, the main effect self-adapting stripping is carried out to the feature data after standardization;The data after stripping is based on residual correlation mining and is handled by interactive effect decoupling;Expert constraint and dynamic benchmark are introduced;Finally output the prediction value of petroleum refining parameter and attribution report.The application adopts layer residual stripping mechanism, and the prediction result is explicitly decomposed into two parts of main effect superposition and interactive effect correction.User can not only obtain prediction value, but also directly see the specific contribution and form of each feature, generate structured attribution report, greatly improve the trust degree.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent modeling and optimization control technology for petrochemical processes, specifically involving a method for predicting and attributing petroleum refining processes based on mechanistic topological hierarchical decoupling. Background Technology

[0002] In modern petroleum refining, the operational status of core reaction units directly determines product yield, quality, and production safety. Such chemical systems typically exhibit complex characteristics including strong multivariate coupling, high nonlinearity, and large time delays. For example, in polymerization reactions, complex interactions exist between physical quantities such as temperature, pressure, catalyst activity, and feed rate; even minute parameter fluctuations can lead to drastic shifts in product performance indicators through nonlinear amplification effects. Therefore, when establishing optimization control and fault early warning models for chemical processes, users not only focus on the model's "predictive accuracy" but also place extreme emphasis on its "logical interpretability" and "physical consistency."

[0003] However, existing modeling techniques have significant limitations in such high-risk scenarios:

[0004] First, traditional statistical modeling, such as linear regression and partial least squares (PLS), typically assumes that variables are independent or that there is only a linear relationship, making it difficult to capture the complex synergistic or antagonistic effects commonly found in chemical reactions. For example, under certain pressures, increasing the temperature can improve the yield, but after exceeding the critical point, the yield drops precipitously due to a sharp increase in side reactions. Linear models cannot fit such non-monotonic characteristics, resulting in insufficient control precision.

[0005] Secondly, while data-driven algorithms based on deep learning, such as DNN and LSTM, offer high accuracy, they are essentially "black box" models lacking transparency. Lacking the constraints of prior physical principles, purely data-driven models are prone to overfitting noise in the data, leading to predictions that violate physical laws or chemical engineering principles. For example, a model might incorrectly predict a false pattern where "output increases when feed is interrupted." Such mechanistic inconsistencies can cause serious safety hazards in real-time control, making it difficult for frontline engineers to trust and adopt such models.

[0006] Furthermore, while some existing interpretable machine learning methods, such as the generalized additive model GA2M, improve model transparency to some extent, their interaction term calculations are often based on the "fully connected" assumption, meaning that all variables are assumed to interact. This ignores the actual physical topology of the chemical plant, such as pipeline connections and material flow constraints. This blind interaction calculation not only leads to an exponential explosion in computational complexity with variable dimensionality but also easily introduces a large number of physically meaningless "spurious correlation interactions," severely interfering with the accuracy of attribution analysis.

[0007] Meanwhile, existing ex-post explanatory tools, such as SHAP and LIME, only provide linear approximations of local predictions and do not analyze the model's intrinsic structure, thus failing to accurately isolate the deep coupling mechanisms between variables. Furthermore, existing attribution methods typically use global static averages as a benchmark, while chemical processes often face dynamic changes such as catalyst activity decay and raw material property fluctuations. Therefore, a dynamic, time-varying benchmark-based attribution analysis mechanism is urgently needed in practical applications to accurately identify the root causes of anomalies under current operating conditions.

[0008] In summary, existing technologies struggle to achieve a balance between "accuracy in complex nonlinear fitting," "logical interpretability consistent with chemical mechanisms," and "dynamic operating condition attribution." There is an urgent need for a novel modeling method that can automatically decouple process parameter interactions, possess both physical topology and mechanistic constraints, and generate structured attribution reports. Summary of the Invention

[0009] Purpose of the invention: The purpose of this invention is to provide a method for predicting and attributing petroleum refining processes based on mechanistic topological hierarchical decoupling.

[0010] To address the issues of lack of physical consistency in existing "black box models" and the proliferation of "spurious correlations" in "general interpretable models" due to neglect of device topology, this invention adopts a layered decoupling mechanism that integrates "spatiotemporal dual mechanism constraints": In the logical dimension, dynamic monotonicity verification is introduced to ensure that the main effect fitting trend conforms to basic chemical engineering principles, thus solving the problem of models "violating common sense"; in the spatial dimension, mechanism topology information is introduced to limit the search space of interaction terms, eliminating false correlations that violate physical connections, thus solving the problem of attribution "calling a deer a horse".

[0011] Ultimately, this invention not only provides high-precision predictions, but also generates structured attribution reports that conform to common chemical engineering knowledge based on dynamic time-varying benchmarks, enabling accurate diagnosis of the operating status of the equipment.

[0012] Technical solution:

[0013] The present invention provides a method for predicting and attributing petroleum refining processes based on mechanistic topological hierarchical decoupling, comprising the following steps:

[0014] Step 1: Collect time-series data and process flow information of the original petroleum refining process, perform multi-dimensional feature cleaning and standardization on the time-series data, and construct a physical adjacency matrix representing the physical connection relationship between variables based on the process flow information to form a standardized feature set.

[0015] Step 2: Based on the dynamic monotonicity constraint mechanism, adaptive decoupling and fitting of the main effects are performed on the standardized feature set to obtain the main effect components. Then, the monotonic trend components of each feature acting independently on the prediction target are extracted, and the first-order residuals are calculated.

[0016] Step 3: Generate a topological mask based on the physical adjacency matrix. First, decouple the first-order residuals from the restricted interaction effects. Then, perform nonlinear interaction fitting between the pairs of variables that are physically connected to obtain the interaction effect components and the second-order residuals that conform to the mechanism.

[0017] Step 4: Construct a dynamic time-varying benchmark based on historical steady-state operating conditions, and combine the main effect components obtained in Step 2 with the interaction effect components obtained in Step 3 to calculate the relative deviation contribution of each feature to the prediction result at the current time.

[0018] Step 5: Finally, output the predicted values ​​of key parameters for petroleum refining and generate a mechanism consistency attribution report that includes a main effect trend diagram and an interaction network diagram.

[0019] Furthermore, step 1 specifically includes the following steps:

[0020] Step 1.1: Collect historical time-series data of the petroleum refining unit and define the refining process state matrix as follows. ,in This represents the number of sampling points for the time series. Define the feature number of process parameters including temperature, pressure, liquid level, and flow rate; define the key performance index vector. , Represents a real number matrix;

[0021] Step 1.2: Calculate the global average operating condition baseline for performance indicators. :

[0022]

[0023] in, For the first The actual index value at each sampling time;

[0024] Construct the initial residual vector This serves as the initial input for subsequent hierarchical decoupling, i.e., the fluctuation amount after removing the steady-state mean; subsequently, to eliminate the interference of different physical dimensions on the model weights, the refining process state matrix is... Each column is Z-score normalized and mapped to the standard normal distribution space:

[0025]

[0026] in, and The first The mean and standard deviation of each process parameter, For the first At the [time]th moment The original detection values ​​of each parameter, These are the standardized eigenvalues;

[0027] Step 1.3: Construct the physical adjacency matrix of the device process flow. Based on the process flow diagram (PFD) and piping and instrumentation diagram (P&ID) of the petroleum refining unit, the physical connection relationships between sensors of various process parameters are analyzed, and a physical adjacency matrix is ​​constructed. :

[0028]

[0029] Here, parameters j and k are two process parameters with different meanings within the same industrial system and at the same time, and the matrix is... This serves as a topological constraint mask for decoupling interaction effects in subsequent steps.

[0030] Furthermore, step 2 specifically includes the following steps:

[0031] Step 2.1: For each feature in the feature set, construct three parallel regressors to compete for fitting:

[0032] Linear Mode: ;

[0033] Nonlinear saturation mode (Poly Mode): ;

[0034] Mono-Tree Mode (Mono-constrained Nonparametric Mode) ;

[0035] in, , These are the coefficients of the linear and quadratic terms, respectively. For input variables, This is the bias value. For decision tree model functions;

[0036] Step 2.2: Calculate the current residual using three modes. coefficient of determination To achieve survival of the fittest, define a decision function. :

[0037]

[0038]

[0039] in, The features are represented by lin, poly, and tree, which correspond to the three parallel regressors in step 2.1, respectively; if the optimal model... Exceeding the preset threshold If so, it is included in the rule base, and the effect is removed from the current residual:

[0040]

[0041] Indicates targeting features The optimal regression model function selected;

[0042] This step is repeated until the independent contributions of all features have been stripped away.

[0043] Furthermore, step 3 specifically includes the following steps:

[0044] Step 3.1: Decoupling based on restricted interaction effects using topological masks, utilizing the physical adjacency matrix. Generate a set of restricted interaction pairs Modeling physically connected pairs of process parameters, and using mechanistic constraints to eliminate spurious associations and compress the search space:

[0045]

[0046] in, The constraint is that the feature indices are not equal. Physical adjacency matrix Topological constraint mask;

[0047] Step 3.2, for the set Each pair of features Build an interaction model For first-order residuals Perform fitting:

[0048]

[0049] in, For the first feature set The and the first One original feature vector, For the first The and the first The interaction terms after feature standardization, The interaction coefficient;

[0050] Step 3.3: Select strong interaction terms through significance testing, and based on... The symbols are given physical meaning:

[0051] like Defined as a synergistic gain effect; if This is defined as an antagonistic interference effect;

[0052] The global average baseline, main effect components, and interaction effect components are layered and superimposed to obtain the final prediction model:

[0053]

[0054] in, This represents the extracted main effects function. The decoupled interaction effect function, For each feature in the feature set, this process is performed using a topological mask, i.e., a conditional mask. It replaces the traditional full permutation search and avoids the curse of dimensionality under high-dimensional sparse data.

[0055] Furthermore, step 4 specifically involves: introducing an expert knowledge base that includes the process operation boundaries and reaction kinetics directions of the refining and chemical unit. and set dynamic reference vector ;

[0056] When inputting new samples First, a consistency check is performed:

[0057]

[0058] in, This is the model's predicted value for the current sample. The target variable is defined as the range of physical constraints based on physical or chemical laws, where... The lower boundary value is the physical constraint range. The upper boundary value of the physical constraint range, For the model to feature The local partial derivative sensitivity, The expected monotonicity direction of the features defined in the expert knowledge base;

[0059] If the constraint is violated, the mechanism safety valve is triggered, introducing a penalty term based on the distance from the constraint boundary. Smoothing or truncation of predicted values:

[0060]

[0061] in, The predicted value is after smoothing or shutting off by the mechanism safety valve. The penalty term is Distance, and the predicted value is Distance. Exceeding physical constraint boundaries Euclidean distance, For the physically feasible region;

[0062] Finally, based on the dynamic benchmark, the marginal contribution of each factor is calculated, and an interpretability report is generated:

[0063]

[0064] Where F represents the pre-trained feature-specific training. The prediction function mapping, As the baseline state, This is the current state.

[0065] The present invention also discloses a computer device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method of the present invention.

[0066] The present invention also discloses a computer-readable storage medium having a computer program / instructions stored thereon, which, when executed by a processor, implements the steps of the method of the present invention.

[0067] The present invention also discloses a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the method of the present invention.

[0068] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages:

[0069] 1. "White-box" transparent prediction of deep fusion mechanism

[0070] Unlike black-box models such as neural networks, which are difficult to interpret, this invention innovatively employs a hierarchical decoupling and residual stripping mechanism to break down complex chemical reaction processes into two intuitive parts: a "main effect baseline" and an "interaction effect correction." Each single-factor response curve generated by the system is forcibly constrained to kinetic monotonicity, ensuring that the model not only predicts accurately but also that its internal logic fully conforms to chemical principles, greatly enhancing the trust of process experts and operators in the AI ​​system.

[0071] 2. Interactive pruning based on physical topology effectively overcomes the curse of dimensionality.

[0072] To address the issues of traditional algorithms, such as GA2M, which employ full permutation search when mining feature interactions, leading to explosive computational costs and spurious correlations, this invention introduces a process mechanism topology mask. It models interactions only for parameter pairs with physical material connections or heat exchange. This restricted search strategy reduces computational complexity from quadratic to linear, completely eliminating statistically significant but physically meaningless "spurious interactions" in the data, significantly improving the model's robustness and generalization ability with small sample sizes.

[0073] 3. Adaptive competitive fitting with a "mechanistic safety valve"

[0074] This invention overcomes the rigidity of traditional mechanistic models, which require manually pre-defined function forms, such as linear or logarithmic ones. The built-in constrained multimodal competition mechanism automatically matches the optimal fitting form (linear, saturated nonlinear, or constrained fluctuation) for each process parameter while satisfying mechanistic constraints. Simultaneously, the mechanistic consistency verification and safety truncation techniques during the inference phase effectively eliminate the risk of model outputs violating physical laws under extreme conditions, ensuring the safety of the control system.

[0075] 4. Dynamic benchmark attribution, enabling real-time operating condition diagnosis.

[0076] Unlike traditional models that only output a single predicted value, this invention establishes a dynamic time-varying benchmark based on historical optimal steady-state conditions. By calculating the deviation contribution of the current operating condition from the dynamic benchmark, the system can quantify the specific impact of each process parameter on performance degradation in real time. This not only achieves prediction but also directly translates into actionable fault diagnosis and operation optimization suggestions, filling the gap in industrial field operations from "data prediction" to "decision execution." Attached Figure Description

[0077] Figure 1 This is a flowchart of the present invention.

[0078] Figure 2 This is a topology injection diagram representing an embodiment of the present invention.

[0079] Figure 3 The bar charts provide visualization for the embodiments of the present invention.

[0080] Figure 4 This is a pie chart illustrating the contribution value of data visualization for this invention.

[0081] Figure 5 This is a comparison chart of the fitted curves of predicted and actual values ​​for the example dataset of this invention. Detailed Implementation

[0082] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0083] The complete workflow of the industrial prediction model establishment and attribution method based on mechanism fusion and hierarchical residual stripping of this invention is as follows: Figure 1 As shown, the system first integrates raw multidimensional monitoring data and process mechanism diagram (P&ID) information in parallel. In the preprocessing stage, in addition to cleaning and standardizing the time series data, the core step is to construct a physical adjacency matrix describing the physical connection relationships between variables based on the P&ID information, and extract the historical steady-state operating conditions of the device as a global benchmark.

[0084] Subsequently, the system enters the adaptive main effect stripping stage. The system initiates a multimodal competition mechanism with embedded monotonicity constraints, and uses basis functions such as linear response, nonlinear saturation, and restricted spline regression to perform parallel fitting of a single feature; the best fit arbitration module automatically determines the physical form of each feature (such as positive correlation linearity, saturated growth, or restricted fluctuation), strips out the main effects, and generates first-order residuals.

[0085] Next, the system enters the interaction mining stage based on topological masks. Instead of blindly traversing all feature combinations, the system uses the aforementioned physical adjacency matrix as a mask to generate candidate interaction terms only for variable pairs with physical connections or heat exchange. Based on the first-order residuals, these candidate interaction terms are fitted and their significance is determined to uncover deep coupling relationships that conform to the mechanism. Finally, an ensemble prediction model containing significant main effects and legitimate interaction effects is trained.

[0086] When a new sample to be predicted is input, the system enters the online prediction and mechanism verification phase. First, the system uses an ensemble model to generate preliminary predictions, which are immediately imported into an expert knowledge base containing physical boundaries and gradient signs for mechanism consistency verification. If the predicted values ​​are found to violate physical common sense, such as yield exceeding limits or gradient reversal, a smoothing correction mechanism based on a penalty term is initiated. Finally, combined with a dynamic time-varying benchmark, the relative deviation contribution of each feature at the current time is calculated, and the corrected high-precision predictions, a structured visual attribution report, and a confidence assessment are output. This process ensures that the model maintains high accuracy while possessing strong interpretability and physical consistency consistent with common sense in the chemical engineering field.

[0087] A complete closed-loop architecture of "data mechanism dual-driven - constrained pattern mining - consistency verification - dynamic attribution" is constructed based on the multivariate complex system prediction and attribution system based on hierarchical residual stripping.

[0088] The system first uses a data sensing module to denoise and standardize the raw, multidimensional, heterogeneous monitoring data from the chemical DCS, including but not limited to reaction temperature, reboiler pressure, feed flow rate, liquid level, and catalyst concentration, and simultaneously analyzes the process topology. Subsequently, the core constrained adaptive mining module, based on the "layered residual stripping" strategy, identifies the main effect form of a single process parameter in parallel under monotonicity constraints, and uses physical topology masks to guide residual iteration, accurately locates the real interactive coupling relationship between process variables, eliminates spurious correlations, and generates a completely transparent white-box model expression.

[0089] Building upon this foundation, the mechanism verification and dynamic inference module incorporates an external expert knowledge base during the model application phase. Through mechanism consistency verification and violation correction mechanisms, it forces the model to output reasonable results that satisfy the material balance and reaction kinetic boundaries. Simultaneously, the system establishes a dynamic time-varying baseline, quantifying the marginal contribution of each process factor relative to the steady-state baseline in real time. Finally, a structured attribution report is generated through a visual interactive front-end, intuitively demonstrating "due to temperature..." High This leads to a lower yield. decline The system employs a causal logic to assist operators in making critical decisions. It achieves a leap from black-box data fitting to transparent mechanism fusion. Its core technology lies in physical topology-based interactive pruning and mechanism consistency verification during the inference stage, ensuring that it can maintain extremely high prediction accuracy and logical interpretability even in complex and ever-changing chemical production environments.

[0090] Example:

[0091] To verify the effectiveness of this invention in handling strongly coupled, nonlinear chemical processes, a typical chemical reactor yield prediction scenario was constructed.

[0092] Set the yield of this chemical reaction. (%) Affected by four measurable process parameters:

[0093] : Reaction temperature (T), unit: °C;

[0094] : Reaction pressure (P), unit: MPa;

[0095] Catalyst flow rate (C), unit: kg / h;

[0096] : Concentration (D) of reactants, in mol / L.

[0097] Based on the real principles of chemical thermodynamics and kinetics, the following generation rules are set:

[0098] Under normal circumstances, within a certain range, there should be an optimal value. To maximize Y, Y should follow , The temperature increases with increasing pressure, and based on the temperature-pressure coupling effect (thermodynamics), there exists an "optimal reaction temperature" that varies with pressure. The higher the pressure, the higher the optimal temperature. When the actual temperature deviates from this optimal point, the yield decreases parabolically.

[0099]

[0100]

[0101] Based on catalyst-concentration antagonism (kinetics): and Increasing either one can improve productivity, but if both are too high at the same time, a "crowding effect" or side effects may occur, leading to diminishing marginal returns.

[0102]

[0103] Final yield:

[0104]

[0105] Based on the above formula, 6000 sets of simulated data were generated, of which 80% were used for training and 20% for testing.

[0106] Step 1: After injecting and initializing the system access data, a simplified process mechanism diagram (P&ID) is simultaneously input. The diagram shows: (Temperature) and (Pressure) belongs to the thermodynamic environment of the same reaction vessel. and This belongs to the material feeding system. The system constructs physical adjacency moments based on this, such as... Figure 2 As shown, the marking and By identifying potential interaction pairs and filtering out physically unrelated combinations (such as the interaction between temperature and catalyst flow rate), the search space is significantly reduced.

[0107] Step 2: Limiting main effects stripping, initiating multimodal competitive fitting of the system. For and The system automatically identifies it as a "linear response mode" (consistent with the principles of dynamics). and Because of the existence of an optimal point, the system identifies it as a "nonlinear saturation mode" and fits it as a parabola and a bell shape, preserving the quadratic term characteristics.

[0108] Step 3: Topology-based interactive mining. Based on the mask generated in Step 1, the system focuses on mining the residuals... and Combining interaction terms for fitting. Successfully captured. and The cooperative offset relationship, i.e., the phenomenon of the optimal temperature shifting with pressure, was successfully captured. and The negative interaction coefficient quantifies the crowding effect.

[0109] The final output pattern is shown in Table 1 below:

[0110] Table 1

[0111]

[0112] And generate a visual contribution bar chart, such as Figure 3 As shown in the figure, the contribution of each factor and interaction message to Y can be clearly seen from the figure.

[0113] Step 4: Mechanism Verification and Prediction. We first perform predictions on the test set using randomly input values. For individual extreme input samples, the system uses an expert constraint module (such as yield cap) to... Automatic truncation correction was performed to ensure that the output results conform to the laws of physics. The test data and results are shown in Table 2 below.

[0114] Table 2

[0115]

[0116] It can be seen that the relative errors are all less than 1%, and the fitting effect is quite good.

[0117] We randomly select a set of data, such as data number 1, and have the system generate a contribution value table as shown in Table 3 below, and generate a visual pie chart of the contribution values ​​as shown below. Figure 4 As shown.

[0118] Table 3

[0119]

[0120] Finally, the predictions were validated on the entire dataset. The fitted curves of the predicted and actual values ​​are shown below. Figure 5 As shown, the model fit is as high as 0.9981, and the prediction results are distributed exactly around the ideal line, which is extremely good.

[0121] For identical datasets, traditional methods were used for analysis, and the results are shown in Table 4 below:

[0122] Table 4

[0123]

[0124] As can be seen from Table 4:

[0125] Regarding accuracy: This system ( This system significantly outperforms traditional linear regression and even surpasses the powerful random forest algorithm. This is because the system accurately characterizes the complex coupling of "temperature-pressure" through explicit interaction terms, while random forests are prone to step-like errors at the boundaries.

[0126] In terms of interpretability: This system not only provides accurate predictions, but also outputs structured attribution reports. This mechanism-based deep attribution is something that traditional algorithms cannot provide.

[0127] In summary, this embodiment verifies that the present invention maintains extremely high prediction accuracy while possessing excellent physical interpretability and logical consistency.

[0128] The above embodiments are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make several improvements and equivalent substitutions without departing from the principle of the present invention. All such improvements and equivalent substitutions to the claims of the present invention fall within the protection scope of the present invention.

Claims

1. A method for predicting and attributing petroleum refining processes based on mechanistic topological hierarchical decoupling, characterized in that, Includes the following steps: Step 1: Collect time-series data and process flow information of the original petroleum refining process, perform multi-dimensional feature cleaning and standardization on the time-series data, and construct a physical adjacency matrix representing the physical connection relationship between variables based on the process flow information to form a standardized feature set. Step 2: Based on the dynamic monotonicity constraint mechanism, adaptive decoupling and fitting of the main effects are performed on the standardized feature set to obtain the main effect components. Then, the monotonic trend components of each feature acting independently on the prediction target are extracted, and the first-order residuals are calculated. Step 3: Generate a topological mask based on the physical adjacency matrix. First, decouple the first-order residuals from the restricted interaction effects. Then, perform nonlinear interaction fitting between the pairs of variables that are physically connected to obtain the interaction effect components and the second-order residuals that conform to the mechanism. Step 4: Construct a dynamic time-varying benchmark based on historical steady-state operating conditions, and combine the main effect components obtained in Step 2 with the interaction effect components obtained in Step 3 to calculate the relative deviation contribution of each feature to the prediction result at the current time. Step 5: Finally, output the predicted values ​​of key parameters for petroleum refining and generate a mechanism consistency attribution report that includes a main effect trend diagram and an interaction network diagram.

2. The petroleum refining prediction and attribution method based on mechanistic topological hierarchical decoupling according to claim 1, characterized in that, Step 1 specifically includes the following steps: Step 1.1: Collect historical time-series data of the petroleum refining unit and define the refining process state matrix as follows. ,in This represents the number of sampling points for the time series. It includes process parameter characteristics such as temperature, pressure, liquid level, and flow rate; Define Key Performance Indicator Vector , Represents a real number matrix; Step 1.2: Calculate the global average operating condition baseline for performance indicators. : ; in, For the first The actual index value at each sampling time; Construct the initial residual vector This serves as the initial input for subsequent hierarchical decoupling, i.e., the fluctuation amount after removing the steady-state mean; subsequently, to eliminate the interference of different physical dimensions on the model weights, the refining process state matrix is... Each column is Z-score normalized and mapped to the standard normal distribution space: ; in, and The first The mean and standard deviation of each process parameter, For the first At the [time]th moment The original detection values ​​of each parameter, These are the standardized eigenvalues; Step 1.3: Construct the physical adjacency matrix of the device process flow. Based on the process flow diagram (PFD) and piping and instrumentation diagram (P&ID) of the petroleum refining unit, the physical connection relationships between sensors of various process parameters are analyzed, and a physical adjacency matrix is ​​constructed. : ; Here, parameters j and k are two process parameters with different meanings within the same industrial system and at the same time, and the matrix is... This serves as a topological constraint mask for decoupling interaction effects in subsequent steps.

3. The petroleum refining prediction and attribution method based on mechanistic topological hierarchical decoupling according to claim 1, characterized in that, Step 2 specifically includes the following steps: Step 2.1: For each feature in the feature set, construct three parallel regressors to compete for fitting: Linear Mode: ; Nonlinear saturation mode (Poly Mode): ; Mono-Tree Mode (Mono-constrained Nonparametric Mode) ; in, , These are the coefficients of the linear and quadratic terms, respectively. For input variables, This is the bias value. For decision tree model functions; Step 2.2: Calculate the current residual using three modes. coefficient of determination To achieve survival of the fittest, define a decision function. : ; ; in, The features are represented by lin, poly, and tree, which correspond to the three parallel regressors in step 2.1, respectively; if the optimal model... Exceeding the preset threshold If so, it is included in the rule base, and the effect is removed from the current residual: ; Indicates targeting features The optimal regression model function selected; This step is repeated until the independent contributions of all features have been stripped away.

4. The petroleum refining prediction and attribution method based on mechanistic topological hierarchical decoupling according to claim 2, characterized in that, Step 3 specifically includes the following steps: Step 3.1: Decoupling based on restricted interaction effects using topological masks, utilizing the physical adjacency matrix. Generate a set of restricted interaction pairs Modeling physically connected pairs of process parameters, and using mechanistic constraints to eliminate spurious associations and compress the search space: ; in, The constraint is that the feature indices are not equal. Physical adjacency matrix Topological constraint mask; Step 3.2, for the set Each pair of features Build an interaction model For first-order residuals Perform fitting: ; in, For the first feature set The and the first One original feature vector, For the first The and the first The interaction terms after feature standardization, The interaction coefficient; Step 3.3: Select strong interaction terms through significance testing, and based on... The symbols are given physical meaning: like Defined as a synergistic gain effect; if This is defined as an antagonistic interference effect; The global average baseline, main effect components, and interaction effect components are layered and superimposed to obtain the final prediction model: ; in, This represents the extracted main effects function. The decoupled interaction effect function, For each feature in the feature set, this process is performed using a topological mask, i.e., a conditional mask. It replaces the traditional full permutation search and avoids the curse of dimensionality under high-dimensional sparse data.

5. The petroleum refining prediction and attribution method based on mechanistic topological hierarchical decoupling according to claim 1, characterized in that, Step 4 specifically involves: introducing an expert knowledge base that includes the process operation boundaries and reaction kinetics directions of the refining and chemical unit. and set dynamic reference vector ; When inputting a new sample First, a consistency check is performed: ; in, This is the model's predicted value for the current sample. The target variable is defined as the range of physical constraints based on physical or chemical laws, where... The lower boundary value is the physical constraint range. The upper boundary value of the physical constraint range, For the model to feature The local partial derivative sensitivity, The expected monotonicity direction of the features defined in the expert knowledge base; If the constraint is violated, the mechanism safety valve is triggered, introducing a penalty term based on the distance from the constraint boundary. Smoothing or truncation of predicted values: ; in, The predicted value is after smoothing or shutting off by the mechanism safety valve. The penalty term is Distance, and the predicted value is Distance. Exceeding physical constraint boundaries Euclidean distance, For the physically feasible region; Finally, based on the dynamic benchmark, the marginal contribution of each factor is calculated, and an interpretability report is generated: ; Where F represents the pre-trained feature-specific training. The prediction function mapping, As the baseline state, This is the current state.

6. A computer device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method of claim 1.

7. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method of claim 1.

8. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method of claim 1.