Causal-related control variable discovery method and system
By constructing shadow manifolds and cross-mapping rules to identify causal-related control variables in catalytic cracking devices, the problem of insufficient causal relationship identification in the prior art is solved, and predictive regulation of flue gas pollutant concentration and stable emissions are achieved.
Patent Information
- Application Number
- CN202410108478.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-25
- Publication Date
- 2025-07-29
AI Technical Summary
The prior art cannot effectively identify the causal-related control variables of the predicted concentration of flue gas pollutants in catalytic cracking devices, resulting in the inability to achieve stable emissions, especially in terms of strong equipment coupling and future numerical predictive regulation.
The causal correlation control variable discovery method is used to construct a shadow manifold through time delay coordinate transformation, combining cross-mapping rules and Pearson correlation coefficients, identify the causal correlation between candidate variables and target variables, construct a chemical process feature matrix and traverse the equipment nodes to determine the causal correlation control variables.
It has realized the predictive regulation of the flue gas pollutant concentration of catalytic cracking devices, improved emission stability, solved the problem of relying solely on end governance, and promoted the effective implementation of joint source and end governance.
Smart Images

Figure CN120386393A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of clean emissions, and specifically to a method for discovering causally related control variables and a system for discovering causally related control variables. Background Art
[0002] As an end-treatment facility, a flue gas scrubbing tower is the last checkpoint before the emission of gaseous pollutants from a fluid catalytic cracking unit. To meet the emission limit requirements for flue gas pollutants from the fluid catalytic cracking unit and avoid the reduction of the unit's load due to excessive flue gas pollutant concentration, which affects economic efficiency, most enterprises have now set up flue gas scrubbing towers. However, due to the uncertainties of raw materials, reaction processes, etc., the fluctuation range of flue gas pollutant concentration is relatively large, which may exceed the treatment capacity of the end-treatment facility, further leading to excessive emissions. At this time, it is necessary to adjust some monitoring parameters of the process to achieve stable and up-to-standard emissions of flue gas pollutants in a way that combines "end + source".
[0003] Generally, the scale of a fluid catalytic cracking unit is relatively large, including many devices such as a riser reactor, a regenerator, an external heat exchanger, etc. Multiple control variables are set on each device. In addition, the reaction mechanism inside the device is complex, and the control variables of each device have a complex non-linear mapping relationship with the flue gas pollutant concentration. Therefore, it is difficult to judge the chemical process control variables related to each pollutant based on experience. To achieve "source" treatment, some scholars combine causal theory to identify the control variables that are causally related to the target variable.
[0004] However, there are two problems in the current research. The first is that the related variable discovery methods based on Granger causality theory and transfer entropy causality theory require that each variable in each system is independent and has no correlation. However, the states of upstream and downstream devices in a chemical process often affect each other, and the control variables in the devices have strong coupling and correlation, making the above two types of methods inapplicable to discovering the causal relationship between variables in a chemical process. Second, the current causal analysis methods often take the monitoring time series data of the target variable and candidate cause variables as input, and cannot identify the causal relationship between the candidate cause variables and the future values of the target variable, and thus cannot support predictive regulation. Aiming at the technical problem that the existing solutions cannot determine the causally related control variables of the predicted value of the flue gas pollutant concentration in a fluid catalytic cracking unit, a new solution for discovering causally related control variables needs to be proposed. Summary of the Invention
[0005] The purpose of the embodiments of the present invention is to provide a method and a system for discovering causally related control variables, so as to at least solve the technical problem that the existing solutions cannot determine the causally related control variables of the predicted value of the flue gas pollutant concentration in a fluid catalytic cracking unit.
[0006] To achieve the above object, a method for discovering causally related control variables is provided in the first aspect of the present invention, which is applied to determining causally related control variables for the predicted concentration of flue gas pollutants in a fluid catalytic cracking unit. The method includes: obtaining a predicted concentration value of a target pollutant as a target variable; performing a time-delay coordinate transformation on the target variable to construct a corresponding shadow manifold; estimating candidate variables based on the shadow manifold and cross-mapping rules to obtain a candidate variable set, and obtaining the Pearson correlation coefficient between each candidate variable and the target variable; obtaining a determination relationship of whether there is a causal relationship between the target variable and each candidate variable based on the magnitude relationship between the Pearson correlation coefficient and a preset causal relationship threshold; and traversing the candidate variable set to obtain all causally related control variables of the target variable.
[0007] Optionally, the obtaining of the predicted concentration value of the target pollutant includes: collecting real-time operating parameters of the fluid catalytic cracking unit; training the real-time operating parameters based on a preset pollutant concentration prediction model to obtain predicted values of the concentrations of various pollutants; and screening out the predicted concentration value of the target pollutant from the predicted values of the concentrations of various pollutants based on the type of the target pollutant; wherein the preset pollutant concentration prediction model is trained based on a machine learning model or a deep learning model through historical operating parameters of the fluid catalytic cracking unit.
[0008] Optionally, the performing of the time-delay coordinate transformation on the target variable to construct a corresponding shadow manifold includes: taking the predicted concentration value of the target pollutant as an input, performing a reconstruction after time-delay transformation to obtain a corresponding shadow manifold, expressed as:
[0009]
[0010] where \(t = 1+(E - 1)\tau,\cdots,L\) represents time; \(E\) is the dimension of the space where the shadow manifold is located; \(\tau\in N\) * represents the time-delay step size; is the target variable at time \(t\).
[0011] Optionally, the estimating of candidate variables based on the shadow manifold and cross-mapping rules to obtain a candidate variable set includes: on the shadow manifold, screening a preset number of nearest neighbor points adjacent to the target variable at time \(t\) in the time dimension; calculating the Euclidean distance between each nearest neighbor point and the target variable respectively, and sorting each nearest neighbor point based on the magnitude relationship of the Euclidean distances to obtain a candidate variable set.
[0012] Optionally, obtaining the Pearson correlation coefficient between each candidate variable and the target variable includes: obtaining the distance ratio between the target variable and each neighboring point based on the sorted neighboring point sequence; calculating the weight of each neighboring point based on the distance ratio between the target variable and each neighboring point; predicting the variables of each neighboring point on the shadow manifold based on the weights of each neighboring point; calculating the Pearson correlation coefficient between the predicted value of the variable of each neighboring point and the target variable as the Pearson correlation coefficient between each candidate variable and the target variable.
[0013] Optionally, the method further includes: traversing the target variable and all data in the data set based on the user-set data set to obtain the corresponding Pearson correlation coefficient.
[0014] Optionally, obtaining the determination relationship of whether there is a causal correlation between the target variable and each candidate variable based on the magnitude relationship between the Pearson correlation coefficient and the preset causal correlation threshold includes: if the Pearson correlation coefficient with the target variable is not less than the preset causal correlation threshold, it is determined that the corresponding candidate variable has a causal correlation with the target variable, and the candidate variable is the causal correlation control variable of the target variable; if the Pearson correlation coefficient with the target variable is less than the preset causal correlation threshold, it is determined that the corresponding candidate variable has no causal correlation with the target variable, and the candidate variable is not the causal correlation control variable of the target variable.
[0015] Optionally, the traversal rule of the candidate variable set includes: constructing a chemical process feature matrix representing the pollutant generation process based on the chemical process feature matrix and the design relationship of the target fluid catalytic cracking unit; determining the index relationship between the chemical process feature matrix and each equipment node; based on the index relationship, starting from the equipment node where the target variable is located, traversing the candidate variables included in the starting node and its connected other equipment nodes one by one to determine whether there is a causal correlation between each candidate variable and the target variable in turn; wherein, if there is no causal correlation between the candidate variable of a certain equipment node and the target variable, the upstream equipment node of the equipment node is no longer traversed.
[0016] Optionally, constructing a chemical process feature matrix representing the pollutant generation process includes: determining the number of equipment nodes based on the fluid catalytic cracking unit design information, one equipment corresponding to one equipment node to obtain the equipment node set; describing the connection relationship between each equipment node in the equipment node set through an adjacency matrix based on the connection relationship between each equipment node to obtain the chemical process feature matrix.
[0017] Optionally, the rows of the chemical process characteristic matrix represent chemical process monitoring variables, and the columns represent sampling moments; determining the index relationship between the chemical process characteristic matrix and each equipment node includes: numbering the chemical process monitoring variables in the sampling order, and grouping the chemical process variables based on the equipment nodes to which the chemical process monitoring variables belong, obtaining chemical process data sets under each equipment node; establishing an index relationship between each equipment node and its corresponding chemical process data set.
[0018] In a second aspect of the present invention, a causal correlation control variable discovery system is provided, which is applied to determining causal correlation control variables for the predicted value of the flue gas pollutant concentration in a fluid catalytic cracking unit. The system includes: a collection unit, configured to obtain the predicted value of the concentration of the target pollutant as the target variable; a transformation unit, configured to perform time-delay coordinate transformation on the target variable to construct a corresponding shadow manifold; a processing unit, configured to estimate candidate variables based on the shadow manifold and cross-mapping rules, obtain a set of candidate variables, and obtain the Pearson correlation coefficient between each candidate variable and the target variable; a relationship determination unit, configured to obtain a determination relationship of whether there is a causal correlation between the target variable and each candidate variable based on the magnitude relationship between the Pearson correlation coefficient and a preset causal correlation relationship threshold;
[0019] A result output unit, configured to traverse the set of candidate variables to obtain all causal correlation control variables of the target variable.
[0020] In a third aspect of the present invention, a computer-readable storage medium is provided. Instructions are stored on the computer-readable storage medium, and when running on a computer, the computer is caused to execute the above-mentioned causal correlation control variable discovery method.
[0021] In a fourth aspect of the present invention, an electronic device is provided. The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor is characterized in that it executes the above-mentioned causal correlation control variable discovery method.
[0022] Through the above technical solutions, the present invention discloses the causal relationship between chemical process monitoring variables and future values of flue gas pollution, discovers the key chemical process monitoring variables affecting future values of flue gas pollutants, solves the problem that stable compliance emissions cannot be achieved only relying on "end-of-pipe" treatment, and promotes the control of flue gas pollutants from "end-of-pipe" treatment to "end-of-pipe + source" treatment.
[0023] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent specific embodiment part. Description of the Drawings
[0024] The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and form a part of the specification. Together with the following specific embodiments, they are used to explain the embodiments of the present invention, but do not constitute a limitation to the embodiments of the present invention. In the drawings:
[0025] Figure 1 is a flowchart of the steps of a method for discovering causally related control variables provided by an embodiment of the present invention;
[0026] Figure 2 is a system structure diagram of a system for discovering causally related control variables provided by an embodiment of the present invention. Specific Embodiments
[0027] The following will describe in detail the specific embodiments of the present invention with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.
[0028] As an end-treatment facility, the flue gas scrubber is the last checkpoint before the emission of gas pollutants from a fluid catalytic cracking unit. To meet the emission limit requirements for flue gas pollutants from the fluid catalytic cracking unit and avoid the unit's load reduction due to excessive flue gas pollutant concentration, which affects economic efficiency, most enterprises have now installed flue gas scrubbers. However, due to the uncertainties in raw materials, reaction processes, etc., the fluctuation range of flue gas pollutant concentration is relatively large, which may exceed the treatment capacity of the end-treatment facility and further lead to excessive emissions. At this time, it is necessary to adjust some monitoring parameters of the process to achieve stable and compliant emissions of flue gas pollutants in a way that combines "end + source".
[0029] Generally, the fluid catalytic cracking unit is large in scale, including many devices such as riser reactors, regenerators, and external heat exchangers. Multiple control variables are set on each device. In addition, the reaction mechanism inside the device is complex, and there are mostly complex non-linear mapping relationships between the control variables of each device and the flue gas pollutant concentration. Therefore, it is difficult to judge the chemical process control variables related to each pollutant based on experience. To achieve "source" treatment, some scholars have combined causal theory to identify the control variables that are causally related to the target variable.
[0030] However, there are two problems in current research. The first is that the relevant variable discovery methods based on Granger causality theory and transfer entropy causality theory require that each variable in each system is independent and has no correlation. However, the states of upstream and downstream equipment in chemical processes often affect each other, and the control variables in the equipment have strong coupling and correlation, making the above two methods inapplicable to discovering the causal relationships between variables in chemical processes. Second, current causal analysis methods often use the monitored time series data of target variables and candidate cause variables as inputs, and cannot identify the causal relationships between candidate cause variables and the future values of target variables, thus unable to support predictive regulation.
[0031] Aiming at the technical problem that the existing solutions cannot determine the causal relevant control variables of the predicted values of flue gas pollutant concentrations in a fluid catalytic cracking unit, the solution of the present invention proposes a method for discovering causal relevant control variables, which is applied to determining the causal relevant control variables of the predicted values of flue gas pollutant concentrations in a fluid catalytic cracking unit. The solution of the present invention predicts the monitored variables of each equipment node by constructing a shadow manifold and a cross-mapping rule to obtain candidate variables of the target variable. There may be a causal relationship between these candidate variables and the target variable. By calculating the Pearson correlation coefficient between the target variable and each candidate variable, it is determined whether there is really a causal relationship between the candidate variable and the target variable. In this way, based on the predicted values of pollutant concentrations, it is possible to predict whether there is a causal relationship between each candidate variable and the predicted value, so as to adjust the corresponding candidate variables based on the prediction results and achieve pre-regulation.
[0032] Figure 1 It is the flowchart of the method for discovering causal relevant control variables provided by an embodiment of the present invention. As Figure 1 shown, the embodiment of the present invention provides a method for discovering causal relevant control variables, and the method includes:
[0033] Step S10: Obtain the predicted value of the concentration of the target pollutant as the target variable.
[0034] Specifically, collect the real-time operating parameters of the fluid catalytic cracking unit; based on a preset pollutant concentration prediction model, train the real-time operating parameters to obtain the predicted values of the concentrations of various pollutants; based on the type of target pollutant, screen out the predicted value of the concentration of the target pollutant from the predicted values of the concentrations of various pollutants; wherein, the preset pollutant concentration prediction model is trained based on a machine learning model or a deep learning model through the historical operating parameters of the fluid catalytic cracking unit.
[0035] In a possible implementation, any regression model, such as a machine learning model, a deep learning model, etc., is used to predict the future values of the flue gas pollutant concentrations in a fluid catalytic cracking unit. Let the flue gas pollutant to be predicted be s, then the prediction sequence of length L at time t can be written as {s} = {s1, s2, ..., s L}. Similarly, for any chemical process monitoring variable x, the monitoring sequence of length L at time t can be written as {x} = {x1, x2, ..., x L}.
[0036] Step S20: Perform a time-delay coordinate transformation on the target variable to construct a corresponding shadow manifold,
[0037] Specifically, taking the predicted value of the concentration of the target pollutant as the input, perform a reconstruction after time-delay transformation to obtain the corresponding shadow manifold, which is expressed as:
[0038]
[0039] where t = 1 + (E - 1)τ, ..., L represents time; E is the dimension of the space where the shadow manifold is located; τ ∈ N * represents the time-delay step size; is the target variable at time t.
[0040] In a possible implementation, perform a time-delay coordinate transformation to construct a shadow manifold M. Taking the predicted flue gas pollutant s as the input, its shadow manifold M s is the reconstruction of s after time-delay transformation, and the mathematical definition of M s is:
[0041]
[0042] where t = 1 + (E - 1)τ, ..., L, E is the dimension of the space where the shadow manifold is located (user-defined), τ ∈ N * represents the time-delay step size, then the shadow manifold M of x x .
[0043] Step S30: Estimate the candidate variables based on the shadow manifold and the cross-mapping rule to obtain a candidate variable set, and obtain the Pearson correlation coefficient between each candidate variable and the target variable.
[0044] Specifically, on the shadow manifold, filter a preset number of neighboring points of the target variable adjacent in the time dimension at time t; calculate the Euclidean distances between each neighboring point and the target variable respectively, and sort each neighboring point based on the magnitude relationship of the Euclidean distances to obtain a candidate variable set. Based on the sorted sequence of neighboring points, obtain the distance ratios between the target variable and each neighboring point; calculate the weights of each neighboring point based on the distance ratios between the target variable and each neighboring point; predict the variables of each neighboring point on the shadow manifold based on the weights of each neighboring point; calculate the Pearson correlation coefficient between the predicted variable value of each neighboring point and the target variable as the Pearson correlation coefficient between each candidate variable and the target variable.
[0045] Embodiment 1:
[0046] In a possible implementation manner, based on cross mapping, the manifold M established in S4 s The manifold M established in S4 x The corresponding point at time t on is estimated. For The cross mapping estimation of can be written as Specifically, it includes:
[0047] Step S301: On the manifold M s search for E + 1 neighboring points adjacent in the time dimension defined .
[0048] Step S302: Arrange the neighboring points obtained in S301 in ascending order according to the same Euclidean distance, and label them as t1,..., t (E+1) .
[0049] Step S303: According to the arrangement obtained in S302, calculate the distance ratio ζ of the i-th neighboring point obtained in S301 to i :
[0050]
[0051] where represents the Euclidean distance between two vectors.
[0052] S304: Use the distance ratio obtained in S303 to calculate the weight ω i corresponding to the i-th neighboring point obtained in S301:
[0053] ω i = ζ i / ∑ζ j j = 1...(E + 1)
[0054] where E is the dimension of the space in which the defined shadow manifold lies.
[0055] Step S305: Use M obtained from S301 s Upper of the nearest neighbor points, written as According to the weights ω calculated in S304 i , estimate M x Upper of the corresponding points The mathematical operation process is as follows:
[0056]
[0057] where is the predicted point in the time dimension of the nearest neighbor points, is the predicted value of.
[0058] Step S306: Calculate the obtained from S305 and the defined Pearson correlation coefficient between the two.
[0059] Preferably, based on the data set set by the user, traverse the target variable and all data in the data set to obtain the corresponding Pearson correlation coefficient.
[0060] Furthermore, increase the data set L, loop and execute step S30,, and record the Pearson correlation coefficient ρ(x|s) calculated in S306 when the loop ends. The maximum value of L is specified by the user.
[0061] Step S40: Based on the magnitude relationship between the Pearson correlation coefficient and the preset causal correlation relationship threshold, obtain a definite relationship on whether there is a causal correlation between the target variable and each candidate variable.
[0062] Specifically, the obtaining of the definite relationship on whether there is a causal correlation between the target variable and each candidate variable based on the magnitude relationship between the Pearson correlation coefficient and the preset causal correlation relationship threshold includes: if the Pearson correlation coefficient with the target variable is not less than the preset causal correlation relationship threshold, it is determined that the corresponding candidate variable has a causal correlation with the target variable, and this candidate variable is the causal correlation control variable of the target variable; if the Pearson correlation coefficient with the target variable is less than the preset causal correlation relationship threshold, it is determined that the corresponding candidate variable has no causal correlation with the target variable, and this candidate variable is not the causal correlation control variable of the target variable.
[0063] In a possible implementation, the predicted value of the flue gas pollutant concentration of the defined fluid catalytic cracking unit is set as the target variable s. A defined chemical process monitoring variable is set as the candidate variable x. Traverse the chemical process monitoring variables included in the nodes of the defined computational graph G according to the rules as the candidate variable x. At the end of the traversal, all causally related variables can be obtained.
[0064] Step S50: Traverse the candidate variable set to obtain all causally related control variables of the target variable.
[0065] Specifically, based on the chemical process feature matrix and the design relationship of the target fluid catalytic cracking unit, a chemical process feature matrix representing the pollutant generation process is constructed; the index relationship between the chemical process feature matrix and each equipment node is determined; based on the index relationship, starting from the equipment node where the target variable is located, traverse the candidate variables included in the starting node and its connected other equipment nodes one by one, and determine whether there is a causal relationship between each candidate variable and the target variable in turn; wherein, if there is no causal relationship between the candidate variable of a certain equipment node and the target variable, the upstream equipment node of this equipment node will no longer be traversed.
[0066] Specifically, the construction of the chemical process feature matrix representing the pollutant generation process includes: determining the number of equipment nodes based on the design information of the fluid catalytic cracking unit, with one construction node corresponding to one equipment node to obtain the equipment node set; based on the connection relationship between each equipment node, describe the connection relationship of each equipment node in the equipment node set through the adjacency matrix to obtain the chemical process feature matrix.
[0067] Further, the determination of the index relationship between the chemical process feature matrix and each equipment node includes: the rows of the chemical process feature matrix represent chemical process monitoring variables, and the columns represent sampling times; number the chemical process monitoring variables according to the sampling order, and group the chemical process variables based on the equipment nodes to which each chemical process monitoring variable belongs to obtain the chemical process data sets under each equipment node; establish the index relationship between each equipment node and its corresponding chemical process data set.
[0068] Example 2:
[0069] In a possible implementation, a chemical process computational graph G is constructed. G=(U, C) is a directed acyclic topological graph used to describe the equipment and its connection relationship in the chemical process, including nodes U and connection relationship C. It includes constructing the nodes U of the G. U={u1, u2,..., u N'}\ is the set of nodes of \(G\), representing all the equipment in the chemical process. One node represents one piece of equipment, where \(N'\) is the number of equipment in the chemical process. It also includes constructing the connection relationship \(C\) of the said \(G\). \(C\) is the set of edges of \(G\), representing all the connection relationships between the equipment in the chemical process. Using the adjacency matrix \(A\in R\) N'×N' As the mathematical description of the connection relationship, the row coordinates represent the starting nodes, the column coordinates represent the ending nodes, and the matrix only contains two elements, 0 and 1, where 1 indicates the existence of a connection relationship and 0 indicates the non-existence of a connection relationship.
[0070] Furthermore, construct the chemical process characteristic matrix \(X\) and its index relationship with the equipment, including constructing the chemical process characteristic matrix \(X\in R\) N×P . Each row of the matrix represents a chemical process monitoring variable (temperature, pressure, liquid level, etc.), written as \(x\), \(N\) represents the total number of chemical process monitoring variables. Each column represents a moment of sampling reading, and the columns of the matrix are arranged in the order of sampling readings, that is, \(P\) represents the length of the time series data formed by the readings at multiple sampling moments. Define \(X\) t \(\in R\) N×i to represent the values of all chemical process monitoring variables at the sampling reading moment \(i\). It also includes constructing the index relationship between the chemical process characteristic matrix \(X\) and the equipment. The chemical process monitoring variables are numbered in sequence, that is, \(X\) can be written as \(X = [x\) 1 , \(x\) 2 , \(\cdots\), \(x\) N T . Since the monitoring points are fixed on a certain piece of equipment in the chemical process, the monitoring variables in \(X\) can be grouped according to the equipment to which they belong. For node \(j\), the set of its corresponding monitoring variables is written as where \(n\) j is the number of monitoring variables included in node \(j\), \(j = 1, 2, \cdots, N'\), and \(N'\) is the number of defined nodes.
[0071] Example 3:
[0072] The traversal rules are as follows:
[0073] 1) Start the search from the node where the target variable \(s\) is located, and traverse one by one the monitoring variables included in the current node and its connected nodes;
[0074] 2) Only the first-order connected upstream nodes of the current node are traversed;
[0075] 3) The variables that have been traversed are not traversed repeatedly;
[0076] 4) When there are no causal-related variables defined in a certain node, the upstream nodes of this node are no longer traversed.
[0077] In the embodiments of the present invention, the solution of the present invention constructs a directed unweighted topological graph to represent the upstream and downstream relationships of equipment in an industrial process, and clusters the monitoring variables according to the equipment, so as to introduce the prior knowledge of the industrial process into causal analysis. By setting traversal rules, the target variable and the candidate monitoring variables participating in the analysis are determined one by one for causal analysis. Using the prediction sequence of the target variable and the observed time series data of the candidate monitoring variables as the input of the convergent cross mapping causal analysis method, it is determined whether the candidate monitoring variable is the causal variable of the target variable. The convergent cross mapping causal analysis method is used for causal analysis, which solves the problem that the Granger causal analysis method and the transfer entropy causal analysis method cannot be applied to the situation where the monitoring variables have coupling and do not satisfy the independence assumption.
[0078] Furthermore, with the increasingly strict pollutant emission standards, the idea of treating catalytic cracking flue gas pollutants relying only on "end-of-pipe" equipment is becoming increasingly difficult to meet the requirements of stable and up-to-standard emissions. However, due to the complexity of chemical processes and the coupling between variables, it is difficult for operators to discover the process monitoring variables related to flue gas pollutant concentrations based on experience. By referring to the output of the model proposed in this patent, operators can quickly locate the relevant process monitoring variables and make adjustments to achieve "source" emission reduction. The injection alkali amount prediction model for the catalytic cracking waste gas desulfurization device proposed in the present invention can be directly used on catalytic cracking units equipped with a DCS system. In addition, the model can be embedded as a component into the "Industrial Internet + Environmental Protection" platform for promotion, and users can also complete the training and cloud deployment of the model by importing training data themselves.
[0079] Example 4:
[0080] A method for discovering causal related control variables based on the predicted value of the NOx concentration in the flue gas of a catalytic cracking unit. In this catalytic cracking system, the main equipment related to flue gas pollutant concentrations includes a riser reactor based on the multi-production of isoparaffin process technology, a dilute-phase tubular regenerator in a charring pot, a flue gas scrubber, and a catalyst external heat exchanger. Twenty chemical process monitoring variables on the four main pieces of equipment are selected as candidate variables. Table 1 lists the serial numbers of all candidate variables and the corresponding monitored contents:
[0081]
[0082] Table 1 Serial numbers of all candidate variables and the corresponding monitored contents
[0083] A total of about 51,000 historical data samples are used to train and test the prediction model. After being arranged in the order of sampling time, they are divided into 5 data sets, named Data Set 1-5, and each data set contains about 10,000 samples.
[0084] Implement the causal correlation control variable discovery method based on the predicted values of flue gas pollutant concentrations in a fluid catalytic cracking unit according to the content of the claims. The specific steps and implementation results are as follows:
[0085] S1 Construct a chemical process computational graph G. G = (U, C) is a directed unweighted topological graph used to describe the equipment and its connection relationships in a chemical process, including nodes U and connection relationships C.
[0086] S1.1 Construct the nodes U of G described in S1. U = {u1, u2,..., u N'} is the set of nodes of G, representing all the equipment in the chemical process. One node represents one piece of equipment, where N' is the number of pieces of equipment in the chemical process.
[0087] S1.2 Construct the connection relationship C of G described in S1. C is the set of edges of G, representing all the connection relationships between the equipment in the chemical process. Use the adjacency matrix A ∈ R N'×N' as the mathematical description of the connection relationship. The row coordinates represent the starting nodes, and the column coordinates represent the ending nodes. The matrix only contains two elements, 0 and 1, where 1 indicates the existence of a connection relationship and 0 indicates the non-existence of a connection relationship.
[0088] The adjacency matrix corresponding to the chemical process system topological graph in Example 4 is established as shown in Table 2:
[0089] Node 1 Node 2 Node 3 Node 4 Node 1 1 1 0 0 Node 2 1 1 1 0 Node 3 0 1 1 0 Node 4 0 1 0 1
[0090] Table 2 Adjacency matrix corresponding to the chemical process system topological graph in Example 4
[0091] S2 Construct a chemical process feature matrix X and its index relationship with the equipment.
[0092] S2.1 Construct a chemical process feature matrix X ∈ R N×P . Each row of the matrix represents a chemical process monitoring variable (temperature, pressure, liquid level, etc.), written as x, and N represents the total number of chemical process monitoring variables. Each column represents a sampling reading moment, and the columns of the matrix are arranged in the order of sampling readings, that is, P represents the length of the time series data formed by multiple sampling moment readings. Define X t ∈ R N×i to represent the values of all chemical process monitoring variables at the sampling reading moment i.
[0093] S2.2 Construct the index relationship between the chemical process feature matrix X and the equipment. The chemical process monitoring variables are numbered in order, that is, X can be written as X = [x 1 , x 2 ,..., x N T Since the monitoring points are fixed on a certain device in the chemical process, the monitoring variables in X can be grouped according to the devices to which they belong. For node j, the set of its corresponding monitoring variables is written as where n j is the number of monitoring variables included in node j, j = 1, 2,..., N', and N' is the number of nodes defined in S1.1.
[0094] The monitoring variables of the chemical process and their corresponding node numbers in the fourth embodiment are shown in Table 3:
[0095]
[0096]
[0097] Table 3 Monitoring variables of the chemical process and their corresponding node numbers in the fourth embodiment
[0098] In the fourth embodiment, x 18 is used as the target variable s for causal analysis, and other variables are used as candidate variables.
[0099] S3 performs flue gas pollution concentration prediction and constructs a prediction sequence. Using a time series graph neural network, the future values of the flue gas pollutant concentration of the fluid catalytic cracking unit are predicted. Let the flue gas pollutant to be predicted be s, then the prediction sequence of length L at time t can be written as {s} = {s1, s2,..., s L}. Similarly, for any monitoring variable x defined by S2.1, the monitoring sequence of length L at time t can be written as {x} = {x1, x2,..., x L}.
[0100] S4 performs time-delay coordinate transformation and constructs the shadow manifold M. Taking s defined by S3 as the input, its shadow manifold M s is the reconstruction of s after time-delay transformation. The mathematical definition of M s is as follows:[[]]
[0101]
[0102] where t = 1 + (E - 1)τ,..., L, E is the dimension of the space where the shadow manifold is located (user-defined), and τ ∈ N * represents the time-delay step size. Similarly, the shadow manifold M x of x defined by S3 can be obtained.
[0103] S5 is based on cross mapping, and uses the manifold M s established in S4 to estimate the corresponding point x at time t on the manifold M established in S4. For The cross - mapping estimation can be written as
[0104] S5.1 Search for E + 1 nearest neighbor points adjacent in the time dimension to those defined in S4 on the manifold M s established in S4 .
[0105] S5.2 Arrange the nearest neighbor points obtained in S5.1 in ascending order of the same Euclidean distance and label them as t1,...,t . (E+1)
[0106] S5.3 According to the arrangement obtained in S5.2, calculate the distance ratio ζ of the i - th nearest neighbor point obtained in S5.1 to i :
[0107]
[0108] where represents the Euclidean distance between two vectors
[0109] S5.4 Use the distance ratio obtained in S5.3 to calculate the weight ω i corresponding to the i - th nearest neighbor point obtained in S5.1
[0110] ω i = ζ i / ∑ζ j j = 1...(E + 1)
[0111] where E is the dimension of the space where the shadow manifold defined in S4 is located
[0112] S5.5 Use the nearest neighbor points on M s obtained in S5.1, written as and the weights ω calculated according to S5.4, to estimate the corresponding point i on M x . The mathematical operation process is as follows :
[0113]
[0114] where is the nearest neighbor point of the predicted point in the time dimension, is 's predicted value
[0115] S5.6 Calculate the obtained in S5.5 and the The Pearson correlation coefficient between the two.
[0116] S6 increases L, and S5 is executed in a loop. When the loop ends, the Pearson correlation coefficient ρ(x|s) calculated by S5.6 is recorded. The maximum value of L is specified by the user.
[0117] In Example 4, the candidate variable x 9 and x 14 of the cross - mapping, where x 9 The scatter plot of the predicted value and the true value gathers near the diagonal, proving that the correlation between the true value and the predicted value is strong; the Pearson correlation coefficient finally converges to about 0.89, proving that x 9 has a strong causal correlation with s. For x 14 The scatter plot of the predicted value and the true value is scattered throughout the coordinate system, proving that the correlation between the true value and the predicted value is weak; the Pearson correlation coefficient finally converges to about 0.14, proving that x 14 has a weak causal correlation with s. This experiment can illustrate that there are differences in the causal correlations between different candidate variables and s.
[0118] S7 sets the threshold Threshold = 0.5 for the causal correlation relationship, and compares the size relationship between ρ(x|s) calculated by S6 and Threshold. If ρ(x|s) ≥ 0.5, then the candidate variable x has a causal correlation with the target variable s; if ρ(x|s) < 0.5, then the candidate variable x has no causal correlation with the target variable s.
[0119] S8 sets the predicted value of the flue gas pollutant concentration of the catalytic cracking unit defined by S3 as the target variable s described in S7. Sets a certain chemical process monitoring variable defined by S2.1 as the candidate variable x described in S7. Traverse the chemical process monitoring variables included in the nodes of the computational graph G defined by S1 according to the rules as the candidate variable x, and execute S3 - S7. At the end of the traversal, all causally related variables can be obtained.
[0120] The traversal rules are as follows:
[0121] 1) Start searching from the node where the target variable s is located, and traverse one by one the monitoring variables included in the current node and its connected nodes;
[0122] 2) Only the first - order connected upstream nodes of the current node are traversed;
[0123] 3) Variables that have been traversed are not traversed repeatedly;
[0124] 4) When there are no causally related variables that meet the definition of S7 in a certain node, the upstream nodes of that node are no longer traversed.
[0125] In Embodiment 4, the specific values of the causal correlation ρ(x|s) of all candidate variables, x 1 , x 2 , x 9 , x 10 , x 11 , x 12 are candidate variables that are manually summarized based on chemical process mechanisms and operation experience and have a causal correlation with s. According to the criterion defined in S7, the x 2 , x 9 , x 11 , x 12 causal correlation variables found by the model are consistent with the results of manual screening, proving that this method can discover correct process monitoring variables.
[0126] The confusion matrix method is used to quantitatively analyze the accuracy of the model, and the definition of accuracy precision is as follows:
[0127]
[0128] The definition of recall is as follows:
[0129]
[0130] The causal correlation ρ(x|s) of each variable with the target variable s in different data sets is colored by a heat map according to the value. Experiments prove that the average accuracy of the model in Data Sets 1-5 is 80.16%, and the recall rate is 70.00%. In addition, the causal correlation ρ(x|s) of the variable changes over time, proving the dynamic nature of this method.
[0131] Keeping the computational graph and eigenvector defined in S1 and S2 unchanged, and keeping the prediction model adopted in S3 unchanged, the Granger causality method is used for the discovery of the causal relationship of flue gas pollutant concentration in Embodiment 4, and a comparative experiment is carried out with the convergent cross mapping method. The average precision and recall rate of the two methods in Data Sets 1-5 are shown in Table 4:
[0132]
[0133] Table 4 Average precision and recall rate of the two methods in Data Sets 1-5
[0134] Since the Granger causality method can only characterize the relative strength of the causal relationship, six variables (consistent with the number of causal correlation variables determined manually) with the highest Granger causality coefficients are selected as the results. It is found by experiments that the causal relationship discovery results of the present patent that introduce the convergent cross mapping method have exceeded both the precision and recall rate of the method that introduces the Granger causality method, proving the superiority of the method.
[0135] Example 5:
[0136] A method for discovering causal-related control variables based on the predicted value of flue gas NOx concentration in a fluid catalytic cracking unit. In this fluid catalytic cracking system, the main equipment related to the concentration of flue gas pollutants includes a riser reactor based on the technology of producing more isoparaffins, a countercurrent bed regenerator, a flue gas scrubber, and a catalyst external heat exchanger. 17 chemical process monitoring variables on the four main pieces of equipment are selected as the inputs of the prediction model. Table 5 lists the serial numbers of all the monitoring variables and the corresponding monitoring contents:
[0137]
[0138] Table 5 Serial numbers of monitoring variables and corresponding monitoring contents in Example 5
[0139] A total of about 63,000 historical data samples are used to train and test the prediction model. After being arranged in the chronological order of sampling, they are divided into 6 data sets, named Data Set 6 - 11, and each data set contains about 10,000 samples.
[0140] Implement the method for discovering causal-related control variables based on the predicted value of flue gas pollutant concentration in a fluid catalytic cracking unit according to the content of the claims. The specific steps and implementation results are as follows:
[0141] S1 Construct a chemical process computational graph G. G=(U, C) is a directed unweighted topological graph used to describe the equipment and its connection relationships in a chemical process, including nodes U and connection relationships C.
[0142] S1.1 Construct the nodes U of G described in S1. U={u1, u2,..., u N'} is the set of nodes of G, representing all the equipment in the chemical process. One node represents one piece of equipment, where N' is the number of equipment in the chemical process.
[0143] S1.2 Construct the connection relationship C of G described in S1. C is the set of edges of G, representing all the connection relationships between the equipment in the chemical process. Use the adjacency matrix A∈R N'×N' as the mathematical description of the connection relationship. The row coordinates represent the starting nodes, and the column coordinates represent the ending nodes. The matrix only contains two elements, 0 and 1, where 1 indicates the existence of a connection relationship and 0 indicates the non-existence of a connection relationship.
[0144] The adjacency matrix corresponding to the chemical process system topological graph in Example 5 is established as follows:
[0145] Node 1 Node 2 Node 3 Node 4 Node 1 1 1 0 0 Node 2 1 1 1 0 Node 3 0 1 1 0 Node 4 0 1 0 1
[0146] Table 6 Adjacency matrix corresponding to the chemical process system topological graph in Example 5
[0147] S2 Construct the chemical process feature matrix X and its index relationship with the equipment.
[0148] S2.1 Construct the chemical process feature matrix X ∈ R N×P . Each row of the matrix represents a chemical process monitoring variable (temperature, pressure, liquid level, etc.), denoted as x, and N represents the total number of chemical process monitoring variables. Each column represents a moment of sampling readings, and the columns of the matrix are arranged in the order of sampling readings, that is, P represents the length of the time series data formed by the readings at multiple sampling moments. Define X t ∈ R N×i to represent the values of all chemical process monitoring variables at the sampling reading moment i.
[0149] S2.2 Construct the index relationship between the chemical process feature matrix X and the equipment. The chemical process monitoring variables are numbered in order, that is, X can be written as X = [x 1 , x 2 ,..., x N . T . Since the monitoring points are fixed on a certain equipment in the chemical process, the monitoring variables in X can be grouped according to the equipment they belong to. For node j, the set of its corresponding monitoring variables is written as where n j is the number of monitoring variables included in node j, and j = 1, 2,..., N', where N' is the number of nodes defined in S1.1.
[0150] The chemical process monitoring variables and their corresponding node numbers in Example 5 are shown in Table 7:
[0151]
[0152]
[0153] Table 7 Chemical process monitoring variables and their corresponding node numbers in Example 5
[0154] In Example 5, take x 15 as the target variable s for causal analysis, and other variables are used as candidate variables.
[0155] S3 Perform flue gas pollution concentration prediction and construct a prediction sequence. Use a time series graph neural network to predict the future values of the flue gas pollutant concentration of the fluid catalytic cracking unit. Let the flue gas pollutant to be predicted be s, then the prediction sequence of length L at time t can be written as {s} = {s1, s2,..., s L}. Similarly, for any chemical process monitoring variable x defined by S2.1, the monitoring sequence of length L at time t can be written as {x} = {x1, x2,..., x L}.
[0156] S4 performs time-delay coordinate transformation to construct the shadow manifold M. With the s defined by S3 as the input, its shadow manifold M s is the reconstruction of s after time-delay transformation, M s is mathematically defined as:
[0157]
[0158] where t = 1+(E - 1)τ,...,L, E is the dimension of the space where the shadow manifold is located (user-defined), and τ ∈ N * represents the time-delay step. Similarly, the shadow manifold M x of x defined by S3 can be obtained.
[0159] S5 is based on cross mapping. Using the manifold M s established in S4 to estimate the corresponding point x at time t on the manifold M established in S4. The cross mapping estimate for can be written as
[0160] S5.1 On the manifold M s established in S4, search for E + 1 nearest neighbor points adjacent to the one defined by S4 in the time dimension.
[0161] S5.2 Arrange the nearest neighbor points obtained in S5.1 in ascending order of the Euclidean distance to the one defined by S4, and label them as t1,...,t (E+1) .
[0162] S5.3 According to the arrangement obtained in S5.2, calculate the distance ratio ζ of the i-th nearest neighbor point obtained in S5.1 to the one i :
[0163]
[0164] where represents the Euclidean distance between two vectors.
[0165] S5.4 Use the distance ratio obtained in S5.3 to calculate the weight ω i corresponding to the i-th nearest neighbor point obtained in S5.1:
[0166] ω i = ζ i / ∑ζ j j = 1...(E + 1)
[0167] where E is the dimension of the space of the shadow manifold defined by S4.
[0168] S5.5 Use M obtained from S5.1 s Up of the neighboring points, written as the weights ω calculated according to S5.4 i , estimate M x Up corresponding point of The mathematical operation process is as follows:
[0169]
[0170] where is the predicted point neighboring points in the time dimension, is predicted value of
[0171] S5.6 Calculate the Pearson correlation coefficient between the obtained from S5.5 and the both defined in S5
[0172] S6 Increase L, loop and execute S5, and record the Pearson correlation coefficient ρ(x|s) calculated by S5.6 when the loop ends. The maximum value of L is specified by the user.
[0173] S7 Set the threshold Threshold = 0.5 for the causal correlation relationship, and compare the size relationship between ρ(x|s) calculated by S6 and Threshold. If ρ(x|s) ≥ 0.5, then the candidate variable x has a causal correlation relationship with the target variable s; if ρ(x|s) < 0.5, then the candidate variable x has no causal correlation relationship with the target variable s.
[0174] S8 Set the predicted value of the flue gas pollutant concentration of the fluid catalytic cracking unit defined in S3 as the target variable s in S7. Set a certain chemical process monitoring variable defined in S2.1 as the candidate variable x in S7. Traverse the chemical process monitoring variables included in the nodes of the computational graph G defined in S1 according to the rules as the candidate variable x, and execute S3 - S7. At the end of the traversal, all causally related variables can be obtained.
[0175] The traversal rules are as follows:
[0176] 1) Start the search from the node where the target variable s is located, and traverse one by one the monitoring variables included in the current node and its connected nodes;
[0177] 2) Only the first-order connected upstream nodes of the current node are traversed;
[0178] 3) Variables that have been traversed are not traversed repeatedly;
[0179] 4) When there are no causally related variables defined by S7 in a certain node, the upstream nodes of that node are no longer traversed.
[0180] In Example 5, x 0 , x 5 , x 7 , x 8 , x 9 , x 10 , x 11 are candidate variables that are manually summarized based on chemical process mechanisms and operation experience and have a causal relationship with s. Variables that meet the specified threshold of S7 are candidate variables that the model discovers to have a causal relationship with s.
[0181] The confusion matrix method is used to quantitatively analyze the accuracy of the model. The definition of accuracy precision is as follows:
[0182]
[0183] The definition of recall is as follows:
[0184]
[0185] The causal relationship ρ(x|s) between each variable and the target variable s in different datasets is colored by heatmap according to the value. Experiments prove that the average accuracy of the model in datasets 6 - 11 is 82.58%, and the recall rate is 66.67%. In addition, the causal relationship ρ(x|s) of the variable changes over time, proving the dynamic nature of this method.
[0186] Keeping the computational graph and feature vectors defined by S1 and S2 unchanged, and keeping the prediction model adopted by S3 unchanged, the Granger causality method is used for the discovery of the causal relationship of the NOx concentration of flue gas pollutants in Example 5, and a comparative experiment is conducted with the convergent cross - mapping method. The average precision and recall rate of the two methods in datasets 6 - 11 are shown in Table 8:
[0187]
[0188] Table 8 Average precision and recall rate of the two methods in datasets 6 - 11
[0189] Since the Granger causality method can only characterize the relative strength of the causal relationship, 7 variables with the highest Granger causality coefficients (consistent with the number of causally related variables judged manually) are selected as the results. It is found by experiments that the causal relationship discovery results of the method introducing the convergent cross - mapping method proposed in this patent have exceeded both the precision and recall rate of the method introducing the Granger causality method, proving the superiority of the method.
[0190] Figure 2 It is the system structure diagram of the causal-related control variable discovery system provided by an embodiment of the present invention. As Figure 2 shown, an embodiment of the present invention provides a causal-related control variable discovery system, which is applied to determine the causal-related control variables of the predicted value of the flue gas pollutant concentration in a fluid catalytic cracking unit. The system includes: a collection unit, configured to obtain the predicted value of the concentration of the target pollutant as the target variable; a transformation unit, configured to perform time-delay coordinate transformation on the target variable to construct a corresponding shadow manifold; a processing unit, configured to estimate candidate variables based on the shadow manifold and cross-mapping rules to obtain a candidate variable set, and obtain the Pearson correlation coefficient between each candidate variable and the target variable; a relationship determination unit, configured to obtain a determination relationship of whether there is a causal relationship between the target variable and each candidate variable based on the magnitude relationship between the Pearson correlation coefficient and a preset causal-related relationship threshold; and a result output unit, configured to traverse the candidate variable set to obtain all the causal-related control variables of the target variable.
[0191] An embodiment of the present invention also provides a computer-readable storage medium, on which instructions are stored, and when running on a computer, the computer is caused to execute the above-mentioned causal-related control variable discovery method.
[0192] An embodiment of the present invention also provides an electronic device, the electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the above-mentioned method for synchronously predicting the multi-pollutant concentration of the flue gas in the fluid catalytic cracking unit is implemented.
[0193] Those skilled in the art can understand that all or part of the steps in the methods of the above embodiments can be completed by instructing relevant hardware through a program. The program is stored in a storage medium, including several instructions for causing a single-chip microcomputer, a chip, or a processor to execute all or part of the steps of the methods of the various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0194] The optional embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the embodiments of the present invention are not limited to the specific details in the above embodiments. Within the scope of the technical concept of the embodiments of the present invention, various simple modifications can be made to the technical solutions of the embodiments of the present invention, and these simple modifications all fall within the protection scope of the embodiments of the present invention. In addition, it should be noted that, in the various specific technical features described in the above specific embodiments, they can be combined in any appropriate manner without conflict. To avoid unnecessary repetition, the embodiments of the present invention will not separately describe various possible combination methods.
[0195] In addition, any combination can be made among various different embodiments of the present invention, as long as it does not violate the idea of the embodiments of the present invention, and it should also be regarded as the content disclosed by the embodiments of the present invention.
Claims
1. A method for discovering causally related control variables, which is applied to determining causally related control variables for the predicted values of flue gas pollutant concentrations in a fluid catalytic cracking unit, and is characterized in that, The method includes: Obtaining a concentration prediction value of a target pollutant as a target variable; Performing time-delay coordinate transformation on the target variable to construct a corresponding shadow manifold; Estimating candidate variables based on the shadow manifold and cross-mapping rules to obtain a set of candidate variables, and obtaining the Pearson correlation coefficient between each candidate variable and the target variable; Based on the magnitude relationship between the Pearson correlation coefficient and a preset causal correlation relationship threshold, obtaining a determination relationship on whether there is a causal correlation between the target variable and each candidate variable; Traversing the set of candidate variables to obtain all causal correlation control variables of the target variable.
2. The method according to claim 1, wherein The obtaining of the concentration prediction value of the target pollutant includes: Collecting real-time operation parameters of a fluid catalytic cracking unit; Training the real-time operation parameters based on a preset pollutant concentration prediction model to obtain prediction values of the concentrations of various pollutants; Based on the type of the target pollutant, screening out the concentration prediction value of the target pollutant from the prediction values of the concentrations of various pollutants; wherein, The preset pollutant concentration prediction model is obtained by training based on a machine learning model or a deep learning model through historical operation parameters of the fluid catalytic cracking unit.
3. The method according to claim 1, wherein The performing of time-delay coordinate transformation on the target variable to construct a corresponding shadow manifold includes: Taking the concentration prediction value of the target pollutant as an input, performing a reconstruction after time-delay transformation to obtain a corresponding shadow manifold, expressed as: where t = 1+(E - 1)τ,..., L, representing time; E is the dimension of the space where the shadow manifold is located; τ ∈ N * represents the time delay step; is the target variable at time t.
4. The method according to claim 1, characterized in that The estimating of candidate variables based on the shadow manifold and cross-mapping rules to obtain a set of candidate variables includes: On the shadow manifold, screening a preset number of nearest neighbor points adjacent to the target variable in the time dimension at time t; Calculating the Euclidean distance between each nearest neighbor point and the target variable respectively, and sorting each nearest neighbor point based on the magnitude relationship of the Euclidean distance to obtain a set of candidate variables.
5. The method according to claim 4, wherein The obtaining of the Pearson correlation coefficient between each candidate variable and the target variable includes: Based on the sorted sequence of nearest neighbor points, obtaining the distance ratio between the target variable and each nearest neighbor point; Calculating the weight of each nearest neighbor point based on the distance ratio between the target variable and each nearest neighbor point; Based on the weight of each nearest neighbor point, predicting the variables of each nearest neighbor point on the shadow manifold; Calculating the Pearson correlation coefficient between the variable prediction value of each nearest neighbor point and the target variable as the Pearson correlation coefficient between each candidate variable and the target variable.
6. The method according to claim 5, wherein The method further includes: Based on a data set set by a user, traversing the target variable and all data in the data set to obtain corresponding Pearson correlation coefficients.
7. The method according to claim 1, characterized in that, The obtaining of the determination relationship on whether there is a causal correlation between the target variable and each candidate variable based on the magnitude relationship between the Pearson correlation coefficient and a preset causal correlation relationship threshold includes: If the Pearson correlation coefficient with the target variable is not less than the preset causal correlation relationship threshold, it is determined that the corresponding candidate variable has a causal correlation with the target variable, and this candidate variable is a causal correlation control variable of the target variable; If the Pearson correlation coefficient between a candidate variable and the target variable is less than the preset causal correlation threshold, it is determined that there is no causal correlation between the corresponding candidate variable and the target variable, and this candidate variable is not a causal correlation control variable of the target variable.
8. The method according to claim 1, wherein The traversal rules of the candidate variable set include: Based on the chemical process feature matrix and the design relationship of the target fluid catalytic cracking unit, construct a chemical process feature matrix representing the pollutant generation process; Determine the index relationship between the chemical process feature matrix and each equipment node; Based on the index relationship, starting from the equipment node where the target variable is located, traverse one by one the candidate variables included in the starting node and other equipment nodes connected to it, and determine whether there is a causal correlation between each candidate variable and the target variable in turn; among them, If there is no causal correlation between the candidate variable of a certain equipment node and the target variable, then the upstream equipment node of this equipment node will no longer be traversed.
9. The method according to claim 8, wherein The construction of the chemical process feature matrix representing the pollutant generation process includes: Based on the design information of the fluid catalytic cracking unit, determine the number of equipment nodes, with one equipment corresponding to one equipment node, to obtain the equipment node set; Based on the connection relationship between each equipment node, describe the connection relationship of each equipment node in the equipment node set through an adjacency matrix to obtain the chemical process feature matrix.
10. The method according to claim 8, characterized in that, The rows of the chemical process feature matrix represent chemical process monitoring variables, and the columns represent sampling times; the determination of the index relationship between the chemical process feature matrix and each equipment node includes: Number the chemical process monitoring variables in the sampling order, and group the chemical process variables based on the equipment nodes to which each chemical process monitoring variable belongs to obtain the chemical process data sets under each equipment node; Establish the index relationship between each equipment node and its corresponding chemical process data set.
11. A causal correlation control variable discovery system, which is applied to the determination of causal correlation control variables for the predicted values of flue gas pollutant concentrations in a fluid catalytic cracking unit, is characterized in that, The system includes: An acquisition unit for obtaining the concentration prediction value of the target pollutant as the target variable; A transformation unit for performing time-delay coordinate transformation on the target variable to construct a corresponding shadow manifold; A processing unit for estimating candidate variables based on the shadow manifold and the cross-mapping rule to obtain a candidate variable set, and obtaining the Pearson correlation coefficient between each candidate variable and the target variable; A relationship determination unit for obtaining the determination relationship of whether there is a causal correlation between the target variable and each candidate variable based on the size relationship between the Pearson correlation coefficient and the preset causal correlation threshold; A result output unit for traversing the candidate variable set to obtain all causal correlation control variables of the target variable.
12. A computer-readable storage medium, characterized in that, Instructions are stored on this computer-readable storage medium, and when running on a computer, they cause the computer to execute the causal correlation control variable discovery method described in any one of claims 1-10.
13. An electronic device, the electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the causal correlation control variable discovery method described in any one of claims 1-10.