Performance causal discovery and fault diagnosis method and system in strip steel hot rolling process
Through cloud-edge and end collaboration technology, a performance-driven gated cyclic stacked autoencoder model is built, which solves the problems of data interconnection and dynamic regulation during the hot rolling process of strip steel, realizes high-precision performance monitoring and fault diagnosis, and ensures the safety and stability of the production process.
Patent Information
- Application Number
- CN202510306227.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-08-26
AI Technical Summary
During the hot rolling process of strip steel, it is difficult for the existing technology to accurately explore the key causal relationships affecting product performance from massive data. The data information has not been efficiently interconnected, and existing systems are difficult to flexibly respond to the complex and changing external environment and dynamic product performance control needs.
The cloud-edge and end-end collaboration technology is adopted, and the edge platform uses data acquisition and preprocessing, and the edge platform builds a performance-driven gated loop stacked autoencoder model, and the cloud platform performs performance causal graph merging and fault diagnosis to achieve global performance monitoring and fault root cause identification.
It improves product performance monitoring accuracy and fault diagnosis accuracy, ensures the safety and stability of the production process, and realizes the sharing and collaborative management of data resources.
Smart Images

Figure CN120540259A_ABST
Abstract
Description
Technical field:
[0001] The present invention belongs to the field of industrial process fault diagnosis, and in particular relates to a performance causal analysis and fault diagnosis method and system for a strip hot rolling process. Background technology:
[0002] Currently, the global manufacturing industry is undergoing a profound transformation centered on digitalization, networking, and intelligentization, entering the era of Industry 4.0. Traditional manufacturing relies on expert experience and manual operations, while intelligent manufacturing utilizes automated, digital, and intelligent technologies. Faced with the growing demand for larger, more complex, and more intelligent production in process industries, the use of advanced process monitoring and fault diagnosis technologies to ensure production safety, ensure product performance stability, improve production efficiency, and reduce production costs has become a research hotspot in the field of complex industrial process control.
[0003] In the context of Industry 4.0 and intelligent manufacturing, the functionality of production equipment throughout the hot strip rolling process is becoming increasingly sophisticated, with increasing levels of automation and complex structures. Subsystems are becoming more closely connected, with multiple devices, operating units, and processes interconnected and distributed across the production line's physical locations. Under the control of a comprehensive automated management and control system, each production line process is interconnected and coupled with multiple systems. Sensory information feedback from the production line is intertwined with equipment control and decision-making. The coupled material, energy, and information flows foster synergies within production processes and across system levels, resulting in heterogeneous data from multiple sources, multiple abnormalities, complex sub-process collaboration, and integrated information management. Although the hot strip rolling process already has a highly integrated and comprehensive management and control system, it still faces numerous challenges and difficulties in analyzing and monitoring key product performance information.
[0004] In the existing technology, the performance monitoring of hot strip rolling still has the following problems: First, the complex product manufacturing process, such as the hot strip rolling process, involves a large number of high-dimensional, nonlinear, and strongly coupled process variables. It is difficult to accurately mine the key causal relationships affecting product performance from massive data and establish a real-time monitoring and diagnosis model; second, from the perspective of the horizontal full-process production process, data information omissions and incomplete data collection still exist in each production process, and the data cannot be efficiently interconnected, resulting in the inability to share and collaboratively analyze key performance information in the production process in real time; third, from the perspective of the vertical control level, although the existing multi-layer information management and control system has a certain degree of standardization in data analysis, management and decision-making processes, its static and predefined operating mode makes it difficult to flexibly respond to complex and changing external environments and dynamic product performance control needs; finally, for the emerging intelligent manufacturing paradigm of cloud-edge-end collaborative interaction, how to integrate cloud computing capabilities and edge real-time data processing capabilities in actual applications to build a performance causal relationship discovery and fault diagnosis system suitable for the hot strip rolling process still needs further exploration and practice. Summary of the invention:
[0005] In order to solve the above problems, the present invention provides a performance causal analysis and fault diagnosis method and system for the hot rolling process of strip steel. It uses cloud-edge-end collaborative technology to provide performance monitoring and fault diagnosis for the hot rolling production process of strip steel, fully explores the performance causal relationship in performance monitoring, and improves the accuracy and real-time performance of fault diagnosis.
[0006] In order to achieve the above objectives, the technical solutions adopted in the embodiments of the present invention are as follows:
[0007] In a first aspect, an embodiment of the present invention provides a method for performance causal discovery and fault diagnosis in a strip hot rolling process, the method comprising the following steps:
[0008] Step S1: On the terminal platform, historical process data of each process of the strip hot rolling production line is obtained and pre-processed;
[0009] Step S2: Filter performance-related process variables from the pre-processed historical process data, and construct a historical multidimensional dataset for each process based on the performance-related process variables. The filtered performance-related process variables and historical multidimensional datasets are then uploaded to the corresponding side platform by process.
[0010] Step S3: On the side platform, based on the performance-related process variables, a performance-driven gated recurrent stacked autoencoder (QGRU-SAE) model is constructed; and the model is trained under performance supervision using a historical multidimensional dataset for the current process.
[0011] Step S4: On the side platform, based on the trained QGRU-SAE model, the historical performance characteristics of each process are obtained, and a performance constraint prediction network for each process is constructed. A loss function is set based on the prediction error of performance-related variables, the prediction error of performance variables, performance correlation, and the degree of input connection sparsity, and training is performed.
[0012] Step S5: Based on the trained performance constraint prediction network, the performance causal relationship between process variables within the process is analyzed, and a process performance causal graph is generated, which is then uploaded to the cloud platform.
[0013] Step S6: On the device-side platform, real-time process data from the current sample point of each process is collected and preprocessed. A real-time dataset for each process is constructed based on performance-related process variables. The real-time dataset is then uploaded to the corresponding edge-side platform by process.
[0014] Step S7: On the edge platform, based on the trained QGRU-SAE model, extract performance features from the real-time data set by process, build a process performance monitoring model, and then upload it to the cloud platform;
[0015] Step S8: On the cloud-side platform, the process performance cause-effect graphs are merged to obtain a global performance cause-effect graph;
[0016] Step S9: On the cloud-side platform, based on the process performance monitoring model, the performance monitoring information of each process is obtained and integrated to obtain global performance monitoring information; based on the global performance monitoring information, it is determined whether there is a performance anomaly; if there is no anomaly, the process returns to step S6; if there is an anomaly, the process proceeds to step S10;
[0017] In step S10, a candidate set of fault-related variables is constructed based on the generalized reconstruction contribution graph to isolate the fault variables; and the fault propagation path is identified and the root cause variable of the fault is located based on the global performance causal graph.
[0018] As a preferred embodiment of the present invention, the end side platform corresponds to each equipment in each process of strip hot rolling; the side platform corresponds to the process; and the cloud side platform corresponds to the global hot rolling production line.
[0019] As a preferred embodiment of the present invention, the preprocessing in step S1 includes:
[0020] Step S11, normalizing the historical process data to eliminate the dimensional differences of multi-source data;
[0021] Step S12: using a multiple interpolation algorithm to fill in the missing data in the normalized historical process data.
[0022] As a preferred embodiment of the present invention, screening performance-related process variables in step S2 specifically includes:
[0023] Step S21, calculating the Pearson correlation coefficient of the variables based on the knowledge of the hot rolling process mechanism and the collected data;
[0024] Step S22 : setting a coefficient threshold, and taking variables corresponding to Pearson coefficients greater than the coefficient threshold as screened performance-related process variables.
[0025] As a preferred embodiment of the present invention, the coefficient threshold is set to the empirical threshold ρ th =0.4.
[0026] As a preferred embodiment of the present invention, step S3 constructs and trains a performance-driven gated recurrent stacked autoencoder QGRU-SAE model, specifically including:
[0027] Assume that the process variable is X=[x1,x2,…,x n ]∈R N×n , the performance variable is Y=[y1,y2,…,y m ]∈R N×m , based on the gated recurrent unit GRU, a performance-supervised QGRU is constructed, and the performance information is used to guide the feature learning of the forget gate and hidden state. The calculation process of the performance-supervised QGRU is as follows:
[0028] r(t)=f r (W r *[h(t-1),X(t)]+U r *Y(t)+b r ) (15)
[0029] z(t)=f z (W z *[h(t-1),X(t)]+U z *Y(t)+b z ) (16)
[0030]
[0031] In formulas (15)-(18), r(t) and z(t) are the values of the reset gate and the update gate at time t, respectively; h(t) are the candidate state vector and output vector of the QGRU unit at time t; f(·) represents the activation function; tanh represents the hyperbolic tangent activation function when updating the candidate state vector; X(t) represents [x1, x2, …, x n ]The value of the variable at time t, Y(t) represents [y1,y2,…,y m ]The value of the variable at time t, X(t) and Y(t) serve as the input of the QGRU unit at time t; {Wr ,b r}、{W z ,b z}、{W h ,b h} are the network parameters of reset gate, update gate and hidden gate respectively; U r 、U z 、U h are the training weights of the performance variables in the reset gate, update gate, and hidden gate respectively; ⊙ represents the element-wise product;
[0032] The QGRU-SAE n neural network composed of n layers of QGRU stacked layer by layer is pre-trained layer by layer to extract deep data features. First, the process variable X and performance variable Y are input into QGRU-SAE 1. In the encoding part, the first layer hidden layer feature h is obtained through QGRU learning. 1 , in the decoding part, use h 1 To reconstruct the process variable X and performance variable Y, we get the reconstructed And the first reconstruction By minimizing Implement pre-training of QGRU-SAE 1. The first reconstruction expression and error are:
[0033]
[0034] In formulas (19)-(20), f(·) represents the activation function of the decoding layer, Network parameters corresponding to the process variable and performance variable decoding layers respectively; is the training parameter set of QGRU-SAE 1;
[0035] The hidden layer features h of QGRU-SAE 1 1 And the performance variable Y is input into QGRU-SAE 2, and the deep hidden layer feature h is obtained by QGRU learning at the encoding layer 2 , in the decoding part, use h 2 To reconstruct the hidden layer features h 1 and performance variable Y, and obtain the reconstructed And the second reconstruction By minimizing Implement pre-training of QGRU-SAE 2; the second reconstruction expression and error are:
[0036]
[0037] In formulas (21) and (22), f(·) represents the activation function of the decoding layer, The network parameters corresponding to the hidden layer features and the performance variable decoding layer respectively; is the training parameter set of QGRU-SAE 2;
[0038] Similarly, the stacked QGRU-SAE n network is pre-trained layer by layer under the supervision of the performance variable Y.
[0039] As a preferred embodiment of the present invention, in step S4, a loss function is set based on the prediction error of the performance-related variable, the prediction error of the performance variable, the performance correlation, and the degree of input connection sparsity. The loss function expression of a single prediction network is as follows:
[0040]
[0041] In formula (25), represents the loss function of the k-th prediction network; Corresponding to the predicted value and true value of the target process variable and performance variable respectively; α and β are the hyperparameters corresponding to the penalty term, which are used to control the performance correlation degree of network features and the level of weight sparsity; Υ corr (z,y) is the Spearman correlation coefficient, which represents the hidden feature z k and performance variable y k the degree of relevance; Indicates that the predicted target is x k The i-th input variable x in the network i The connection weight of
[0042] For the performance constraint prediction network, since a single prediction network can only obtain the potential dependent variable set of a single target variable, in order to obtain the overall performance causal skeleton of all process variables, the global loss function is minimized. To train n prediction networks simultaneously, The expression is:
[0043]
[0044] As a preferred embodiment of the present invention, generating a process performance cause-and-effect diagram in step S5 specifically includes:
[0045] Step S51, using the performance constraint prediction network to extract the causal relationship skeleton to obtain a performance causal undirected graph;
[0046] Step S52: using the complexity of the prediction model to determine the causal direction and obtain a performance causal directed graph;
[0047] Step S53: prune the performance causal directed graph to obtain a final performance causal graph.
[0048] In a second aspect, an embodiment of the present invention further provides a performance causal discovery and fault diagnosis device for a strip hot rolling process, the device comprising: a data acquisition module 101, a variable screening module 102, a QGRU-SAE model construction module 103, a prediction network construction module 104, a performance causal skeleton search module 105, a performance causal discovery module 106, a process performance monitoring module 107, a global performance monitoring module 108, and a fault tracing module 109; wherein,
[0049] The data acquisition module 101 is used to acquire historical process data of each process of the strip hot rolling production line on the terminal side platform and perform preprocessing; it is also used to collect real-time process data of the current sample point of each process and perform preprocessing;
[0050] The variable screening module 102 is used to screen performance-related process variables from pre-processed historical process data, and construct a historical multidimensional dataset for each process step based on the performance-related process variables, and then upload the screened performance-related process variables and historical multidimensional dataset to the corresponding side platform by process step. It is also used to construct a real-time dataset for each process step based on the performance-related process variables, and then upload the real-time dataset to the corresponding side platform by process step.
[0051] The QGRU-SAE model construction module 103 is used to construct a performance-driven gated recurrent stacked autoencoder QGRU-SAE model based on the performance-related process variables on the side platform; and use the historical multidimensional dataset of the current process to perform model training under performance supervision;
[0052] The prediction network construction module 104 is used to obtain the historical performance characteristics of each process based on the trained QGRU-SAE model on the side platform, and to construct a performance constraint prediction network for each process. The loss function is set based on the prediction error of performance-related variables, the prediction error of performance variables, the performance correlation, and the degree of input connection sparsity, and the training is performed.
[0053] The performance causal skeleton search module 105 is used to analyze the performance causal relationship between process variables within the process based on the trained performance constraint prediction network, and generate a process performance causal graph, which is then uploaded to the cloud platform;
[0054] The performance causal discovery module 106 is used to merge the process performance causal graphs on the cloud platform to obtain a global performance causal graph;
[0055] The process performance monitoring module 107 is used to extract performance features from the real-time data set by process based on the trained QGRU-SAE model on the edge platform, and build a process performance monitoring model, which is then uploaded to the cloud platform.
[0056] The global performance monitoring module 108 is used to obtain and integrate the performance monitoring information of each process on the cloud-side platform based on the process performance monitoring model to obtain global performance monitoring information; based on the global performance monitoring information, determine whether there is a performance anomaly; if there is no anomaly, start the data acquisition module 101; if there is an anomaly, start the fault tracing module 109;
[0057] The fault tracing module 109 is used to construct a candidate set of fault-related variables based on the generalized reconstruction contribution (GRBC) graph to isolate fault variables; and to identify fault propagation paths and locate fault root variables based on the global performance causal graph.
[0058] In a third aspect, an embodiment of the present invention further provides a performance causal discovery and fault diagnosis system for a hot rolling process of a steel strip, the system comprising an end-side platform, an edge-side platform, and a cloud-side platform; wherein the end-side platform corresponds to specific equipment in each process of the hot rolling production line, the edge-side platform corresponds to each process, and the cloud-side platform corresponds to the entire hot rolling production line;
[0059] The end-side platform is used to obtain historical process data of each process of the strip hot rolling production line and perform preprocessing; screen performance-related process variables from the preprocessed historical process data, and build a historical multidimensional data set for each process based on the performance-related process variables, and then upload the screened performance-related process variables and historical multidimensional data sets to the corresponding side-side platform by process; it is also used to collect real-time process data of the current sample point of each process and perform preprocessing; and build a real-time data set for each process based on the performance-related process variables, and then upload the real-time data set to the corresponding side-side platform by process;
[0060] The side platform is used to build a performance-driven gated recurrent stacked autoencoder QGRU-SAE model based on the performance-related process variables; and use the historical multidimensional data set under the current process to train the model under performance supervision; then based on the trained QGRU-SAE model, the historical performance characteristics of each process are obtained, and a performance constraint prediction network for each process is constructed, and a loss function is set based on the prediction error of the performance-related variables, the prediction error of the performance variables, the performance correlation and the degree of input connection sparsity, and training is performed; based on the trained performance constraint prediction network, the performance causal relationship between the process variables in the process is analyzed, and a process performance causal graph is generated, and then uploaded to the cloud-side platform; and based on the trained QGRU-SAE model, the performance characteristics in the real-time data set are extracted for each process, and a process performance monitoring model is constructed, and then uploaded to the cloud-side platform;
[0061] The cloud-side platform is used to merge the process performance cause-and-effect graphs to obtain a global performance cause-and-effect graph; obtain and fuse the performance monitoring information of each process based on the process performance monitoring model to obtain global performance monitoring information; determine whether there is a performance anomaly based on the global performance monitoring information; if there is no anomaly, proceed to the next sampling point; if there is an anomaly, construct a fault-related variable candidate set based on the generalized reconstruction contribution (GRBC) graph to achieve fault variable isolation; and identify the fault propagation path and locate the fault root variable based on the global performance cause-and-effect graph.
[0062] The performance causal analysis and fault diagnosis method and system for the hot rolling process of steel strip provided by the embodiments of the present invention have the following beneficial effects:
[0063] (1) Design a data management platform to realize the sharing and exchange of data resources on the end and cloud sides. Collect multi-dimensional real-time data of the hot rolling production line on the end side, and complete the integration and storage of data throughout the product life cycle on the cloud side. This can effectively eliminate the information omission and incompleteness caused by the independent collection of data from each process in the traditional model.
[0064] (2) Constructing a performance-driven gated recurrent stacked autoencoder model can fully utilize the performance information of the hot rolling production process, explore the deep-level performance characterization of the production line, and realize the causal relationship mining and dynamic performance monitoring of product performance in the hot rolling process, which greatly improves the accuracy of product performance monitoring and can reliably ensure the safe and stable operation of the production line.
[0065] (3) Based on the cloud-edge-end collaborative architecture, by integrating the data resources and model algorithms on the end side, edge side and cloud side, the update and collaboration of the models between platforms in the hot rolling production process are realized.
[0066] Of course, it is not necessary to achieve all of the advantages described above simultaneously in order to implement any product or method of the present invention. Description of the drawings:
[0067] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0068] Figure 1 This is a flow chart of a method for performance causal discovery and fault diagnosis in a strip hot rolling process provided by an embodiment of the present invention;
[0069] Figure 2 This is a schematic diagram of multi-database collaborative storage on a terminal platform provided by an embodiment of the present invention;
[0070] Figure 3Schematic diagram of the gated recurrent unit in an embodiment of the present invention;
[0071] Figure 4 Schematic diagram of the stacked autoencoder principle in an embodiment of the present invention;
[0072] Figure 5 is a schematic diagram of a prediction network in an embodiment of the present invention;
[0073] Figure 6 2. Schematic diagram of causal skeleton score threshold division in an embodiment of the present invention;
[0074] Figure 7 This is a schematic diagram of a cloud-edge-device collaborative platform in an embodiment of the present invention;
[0075] Figure 8 1 is a schematic structural diagram of a performance causal analysis and fault diagnosis system for a strip hot rolling process according to an embodiment of the present invention. Specific implementation method:
[0076] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations. It should be noted that the embodiments of the present invention and the features in the embodiments can also be combined with each other in the absence of conflict.
[0077] It should be noted that similar reference numerals and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined or explained in subsequent figures. In the description of the present invention, the terms "first," "second," "third," "fourth," etc. are used only to distinguish the description and are not to be understood as indicating or implying relative importance.
[0078] In response to the problems existing in the manufacturing process of complex products, such as the hot rolling process of strip steel, and the demand for exploring rich information resources in production line data under the trend of digital transformation and intelligent upgrading of the manufacturing industry, the embodiments of the present invention provide a cloud-edge-end collaborative strip hot rolling performance monitoring and fault diagnosis method and system, including: on the end-side platform, real-time data of the production line is collected and non-ideal data is filled and performance-related variables are screened; on the edge-side platform, a performance-driven gated recurrent stacked autoencoder model is constructed to mine the deep performance characterization of the multi-dimensional data of the production line; a performance-constrained prediction network is constructed to parse the performance causal relationship between variables, and a deep performance characterization model is used to realize real-time performance monitoring of each process; on the cloud-side platform, upper-level information such as process technology knowledge is fully considered, the causal relationship learned from each process is merged to obtain a global performance causal graph, and the monitoring information of each process is integrated based on Bayesian reasoning to realize dynamic monitoring of global performance; a candidate set of fault variables is constructed based on the generalized reconstruction contribution graph, and the fault propagation path and the root cause variable of the fault are identified in combination with the performance causal graph.
[0079] like Figure 1 As shown, the performance causal discovery and fault diagnosis method of the strip hot rolling process includes the following steps:
[0080] Step S1: On the terminal platform, historical process data of each process of the strip hot rolling production line is obtained and pre-processed;
[0081] Step S2: Filter performance-related process variables from the pre-processed historical process data, and construct a historical multidimensional dataset for each process based on the performance-related process variables. The filtered performance-related process variables and historical multidimensional datasets are then uploaded to the corresponding side platform by process.
[0082] Step S3: On the side platform, a performance-driven gated recurrent stacked autoencoder (QGRU-SAE) model is constructed based on the performance-related process variables; and the model is trained under performance supervision using a historical multidimensional dataset of the current process.
[0083] Step S4: On the side platform, based on the trained QGRU-SAE model, the historical performance characteristics of each process are obtained, and a performance constraint prediction network for each process is constructed. A loss function is set based on the prediction error of performance-related variables, the prediction error of performance variables, performance correlation, and the degree of input connection sparsity, and training is performed.
[0084] Step S5: Based on the trained performance constraint prediction network, the performance causal relationship between process variables within the process is analyzed, and a process performance causal graph is generated, which is then uploaded to the cloud platform.
[0085] Step S6: On the device-side platform, real-time process data from the current sample point of each process is collected and preprocessed. A real-time dataset for each process is constructed based on performance-related process variables. The real-time dataset is then uploaded to the corresponding edge-side platform by process.
[0086] Step S7: On the edge platform, based on the trained QGRU-SAE model, extract performance features from the real-time data set by process, build a process performance monitoring model, and then upload it to the cloud platform;
[0087] Step S8: On the cloud-side platform, the process performance cause-effect graphs are merged to obtain a global performance cause-effect graph;
[0088] Step S9: On the cloud-side platform, based on the process performance monitoring model, the performance monitoring information of each process is obtained and integrated to obtain global performance monitoring information; based on the global performance monitoring information, it is determined whether there is a performance anomaly; if there is no anomaly, the process returns to step S6; if there is an anomaly, the process proceeds to step S10;
[0089] Step S10: construct a candidate set of fault-related variables based on the generalized reconstruction contribution (GRBC) graph to isolate fault variables; and identify the fault propagation path and locate the fault root variable based on the global performance causal graph.
[0090] The method steps are further described below with reference to specific embodiments.
[0091] Step S1: On the end-side platform, historical process data of each process of the strip hot rolling production line is obtained and pre-processed.
[0092] In this step, the end-side platform corresponds to each device in each process of hot rolling of strip steel, and each device corresponds to an end-side platform for data collection. At the same time, when collecting multi-dimensional real-time data of the hot rolling production line of strip steel, relying on the diversified sensor devices and high-performance programmable logic controllers and other industrial-grade hardware infrastructure equipped with the hot rolling production line, through standardized industrial communication protocols and network architectures such as fieldbus and industrial Ethernet, the real-time collection and integration of the operation data of the equipment in each process of the entire production process is realized and stored on the end-side platform. Figure 2 As shown, the end-side platform's database structure consists of three layers: the sensor layer, the pSpace real-time historical database layer, and the multi-class database layer. Based on sensor data from the sensor layer, with the pSpace real-time historical database at its core, a multi-database collaborative storage architecture is constructed, combining the time-series database InfluxDB and the relational database MySQL in the multi-class database layer. This provides a unified data access interface and storage support for cloud-edge-end collaborative computing. Correspondingly, on the cloud side, a data alignment algorithm based on product length percentiles is used to perform spatiotemporal registration of data collected from the same strip product at each production line step, ensuring data consistency and continuity.
[0093] The preprocessing is to address the performance issues of low utilization and many missing values in the raw data collected from the hot rolling production line, in order to ensure the accuracy and reliability of subsequent modeling and analysis. It specifically includes:
[0094] In step S11, the historical process data is normalized to eliminate the dimensional differences of multi-source data. The z-score normalization method is used to convert the data into a standard normal distribution data set with a mean of 0 and a standard deviation of 1 through linear transformation, thereby achieving data standardization.
[0095] In this step, the normalization formula is as follows:
[0096]
[0097] In formula (1), X * , μ, and σ are the normalized value, mean, and variance of the variable X, respectively.
[0098] In step S12, a multiple imputation (MI) algorithm is used to fill in missing data in the normalized real-time data. Based on the Bayesian statistical theory framework, the MI algorithm constructs a probability distribution model for missing data and generates multiple reasonable sets of interpolation values, thereby creating multiple complete data sets. Compared to single interpolation methods, the core advantage of the MI algorithm lies in its ability to fully account for the uncertainty of missing data and effectively reduce interpolation bias by integrating and analyzing multiple sets of interpolation results.
[0099] Specifically, based on the statistical characteristics of the observed data, the Markov Chain Monte Carlo (MCMC) method is used to simulate the posterior distribution of missing values:
[0100]
[0101] In formula (2), X obs and X mis are observed data and missing data respectively, P(θ|X obs ,X mis ) means given X=[X obs ,X mis ] T The conditional density of θ under θ; P(X mis |X obs ) indicates that given X obs Next X mis The conditional density of the posterior distribution P(θ|X obs ) can be calculated by the following formula:
[0102]
[0103] In formula (3), the posterior distribution P(θ|X obs ) can be calculated iteratively: First, for X mis Select a reasonable initial value;
[0104] Secondly, according to the posterior distribution P(θ|X obs ,X mis ) estimate θ;
[0105] Finally, according to the current parameter θ i , in the i+1th iteration we get At the conditional density In the simulation of the i+1th iteration, we get The iterative process is repeated until a Markov chain is generated, which is used to calculate and generate an imputed data set to ensure the reliability of missing value filling and the accuracy of statistical inference.
[0106] Step S2: Filter performance-related process variables from the preprocessed historical process data, and construct a historical multidimensional dataset for each process based on the performance-related process variables. Then, upload the filtered performance-related process variables and historical multidimensional datasets to the corresponding side platform by process.
[0107] In this step, the side platform corresponds to each process, and each process corresponds to a side platform; the end-side platforms of all equipment under each process upload the collected data to the side platform of the corresponding process.
[0108] The screening of performance-related variables takes into account the characteristics of complex product manufacturing processes, such as system-related coupling, process interaction, and hierarchical collaboration. The pre-processed production data is screened according to the degree of performance relevance to reduce data storage costs and avoid model overfitting and reduced accuracy. Specifically, it includes:
[0109] Step S21 , based on the hot rolling process mechanism model and production mode, calculate the Pearson correlation coefficient of the variables.
[0110] In this step, the calculation formula of the Pearson correlation coefficient is:
[0111]
[0112] In formula (4), cov(X,Y) represents the covariance of X and Y, σ X , σ Y They represent the variance of X and Y respectively, and E(·) represents the mean.
[0113] Step S22 : setting a coefficient threshold, and taking variables corresponding to Pearson coefficients greater than the coefficient threshold as screened performance-related process variables.
[0114] In this step, preferably, the coefficient threshold is set to the empirical threshold ρ th =0.4,|ρ X,Y |≥ρ th The variables are screened as performance-related variables, and the screened variables are used to construct a performance-related data set.
[0115] In step S3, a performance-driven gated recurrent stacked autoencoder (QGRU-SAE) model is constructed on the side platform based on the performance-related process variables; and the model is trained under performance supervision using a historical multidimensional dataset of the current process.
[0116] In this step, the construction and training process of the performance-driven gated recurrent stacked autoencoder model includes:
[0117] Considering that complex product manufacturing processes, such as the hot-rolling process of steel strip, are often nonlinear and dynamic, and that material, energy, and information flows interact and interact between various processes, the production line data collected in step S1 contains time-dependent information related to product performance. This paper combines the gated recurrent unit (GRU) and the stacked autoencoder (SAE) to design a performance-supervised feature extraction strategy. Performance-related process variables are process-related parameters and are also process variables.
[0118] Assume that the process variable is X=[x1,x2,…,x n ]∈R N×n , the performance variable is Y=[y1,y2,…,y m ]∈R N×m ,like Figure 3 As shown, the calculation steps of the GRU unit are as follows:
[0119] r(t)=f r (W r *[h(t-1),X(t)]+b r ) (5)
[0120] z(t)=f z (W z *[h(t-1),X(t)]+b z ) (6)
[0121]
[0122] In equations (5)-(8), r(t) and z(t) are the values of the reset gate and the update gate at time t, respectively; h(t) are the candidate state vector and output vector of the GRU unit at time t; f(·) represents the activation function; tanh represents the hyperbolic tangent activation function when updating the candidate state vector; X(t) represents [x1, x2, …, x n ]The value of the variable at time t is the input of the GRU unit at time t; {W r ,b r}、{W z ,b z}、{W h ,b h} are the network parameters of the reset gate, update gate, and hidden gate respectively; ⊙ represents element-wise product.
[0123] SAE is a neural network method composed of multiple autoencoders (AE) stacked layer by layer, such as Figure 4 As shown in Figure 1, SAE achieves the purpose of extracting deep data features through layer-by-layer network pre-training. The process variable X is input into the first layer AE network (denoted as AE 1) and the reconstructed process variable is obtained through encoding and decoding. By minimizing the reconstruction loss To implement the pre-training of AE 1 network, the process is expressed as follows:
[0124]
[0125] In formulas (9)-(11), h 1 is the hidden layer feature extracted by AE 1; f(·) represents the activation function; The network parameters corresponding to the AE 1 encoding layer and decoding layer respectively; is the training parameter set of AE 1. Then the hidden layer feature h of AE 1 is reconstructed in AE 2 1 To extract deeper features h 2 , the encoding-decoding and reconstruction error of AE 2 can be expressed as:
[0126]
[0127] In formulas (12)-(14), h 2 Hidden layer features extracted for AE 2; is the reconstructed output of AE 2; f(·) represents the activation function; The network parameters corresponding to the AE 2 encoding layer and decoding layer respectively; is the training parameter set of AE 2.
[0128] By analogy, SAE can be pre-trained layer by layer until AE n. This network structure enables SAE to reduce redundant information layer by layer and stably extract deep hidden features while fully representing the input data.
[0129] The QGRU-SAE model combines the technical features of GRU and SAE, using QGRU as the encoding layer in each layer of the SAE network structure and extracting deep performance characteristics of the input data in the form of SAE stacking. Performance information Y is added to the QGRU unit as supervision, and the performance information is used to guide the feature learning of the forget gate and hidden state. The performance-supervised QGRU can be expressed as follows:
[0130] Assume that the process variable is X=[x1,x2,…,x n ]∈R N×n , the performance variable is Y=[y1,y2,…,y m ]∈R N×m , based on the gated recurrent unit GRU, a performance-supervised QGRU is constructed, and the performance information is used to guide the feature learning of the forget gate and hidden state. The calculation process of the performance-supervised QGRU is as follows:
[0131] r(t)=f r (W r *[h(t-1),X(t)]+U r *Y(t)+b r ) (15)
[0132] z(t)=f z (W z *[h(t-1),X(t)]+U z *Y(t)+b z ) (16)
[0133]
[0134] In formulas (15)-(18), r(t) and z(t) are the values of the reset gate and the update gate at time t, respectively; h(t) are the candidate state vector and output vector of the QGRU unit at time t; f(·) represents the activation function; tanh represents the hyperbolic tangent activation function when updating the candidate state vector; X(t) represents [x1, x2, …, x n ]The value of the variable at time t, Y(t) represents [y1,y2,…,y m ]The value of the variable at time t, X(t) and Y(t) serve as the input of the QGRU unit at time t; {W r ,b r}、{W z ,b z}、{Wh ,b h} are the network parameters of reset gate, update gate and hidden gate respectively; U r 、U z 、U h are the training weights of the performance variables in the reset gate, update gate, and hidden gate respectively; ⊙ represents the element-wise product;
[0135] The QGRU-SAE n neural network composed of n layers of QGRU stacked layer by layer is pre-trained layer by layer to extract deep data features. First, the process variable X and performance variable Y are input into QGRU-SAE 1. In the encoding part, the first layer hidden layer feature h is obtained through QGRU learning. 1 , in the decoding part, use h 1 To reconstruct the process variable X and performance variable Y, we get the reconstructed And the first reconstruction By minimizing Implement pre-training of QGRU-SAE 1. The first reconstruction expression and error are:
[0136]
[0137] In formulas (19)-(20), f(·) represents the activation function of the decoding layer, Network parameters corresponding to the process variable and performance variable decoding layers respectively; is the training parameter set of QGRU-SAE 1.
[0138] The hidden layer features h of QGRU-SAE 1 1 And the performance variable Y is input into QGRU-SAE 2, and the deep hidden layer feature h is obtained by QGRU learning at the encoding layer 2 , in the decoding part, use h 2 To reconstruct the hidden layer features h 1 and performance variable Y, and obtain the reconstructed And the second reconstruction By minimizing Implement pre-training of QGRU-SAE 2; the second reconstruction expression and error are:
[0139]
[0140] In formulas (21) and (22), f(·) represents the activation function of the decoding layer, The network parameters corresponding to the hidden layer features and the performance variable decoding layer respectively; is the training parameter set of QGRU-SAE 2;
[0141] The stacked QGRU-SAE s network is pre-trained layer by layer under the supervision of the performance variable Y. During the training process, performance-irrelevant information will be gradually discarded. After training, the network can extract the deep dynamic performance characteristics contained in the input data, enhancing the reliability of the performance prediction network and the interpretability of the performance causal graph.
[0142] In step S4, on the side platform, the historical performance characteristics of each process are obtained based on the trained QGRU-SAE model, and a performance constraint prediction network for each process is constructed. The loss function is set based on the prediction error of performance-related variables, the prediction error of performance variables, performance correlation, and the degree of input connection sparsity, and training is performed.
[0143] Specifically, in this step, when setting the loss function of the performance constraint prediction network, a prediction network with the same structure is constructed for each variable in the process. The remaining variables are used as prediction inputs. The Spearman correlation coefficient is used to quantify the correlation between the network's hidden layer features and the performance variables, and this is maximized during the training process. In addition, the input weight is added as a penalty term to the loss function to reduce redundant connections that are irrelevant to performance, and to achieve better target predictions with as few input variables as possible. The specific process includes:
[0144] like Figure 5 As shown, the performance-constrained prediction network is composed of a fully connected neural network, and the input layer integrates the current data feature h k- and the deep dynamic features h extracted from historical data n , the target variable X and performance variable Y are predicted simultaneously through a single hidden layer neural network. In order to make the hidden layer network obtain the most relevant feature information of performance, the Spearman correlation coefficient Y is used. corr (z,y) quantifies the correlation between the hidden feature z and the performance variable Y and maximizes it during training. corr The calculation expression of (z,y) is as follows:
[0145]
[0146] c j =rnk(z j )-rnk(y j ) (twenty four)
[0147] In formulas (23) and (24), rnk(z j ) and rnk(y j ) is the rank of the sample point in the corresponding variable.
[0148] The purpose of constructing the performance-constrained prediction network is to extract the performance causal connection between the input variables and the target variables from the prediction network. In order to reduce redundant connections irrelevant to performance and achieve better target prediction with as few input variables as possible, the prediction network input layer weights are It is added to the loss function as a penalty term to achieve the purpose of sparse input connections.
[0149] The loss function expression of a single prediction network is as follows:
[0150]
[0151] In formula (25), represents the loss function of the k-th prediction network; Corresponding to the predicted value and true value of the target process variable and performance variable respectively; α and β are the hyperparameters corresponding to the penalty term, which are used to control the performance correlation degree of network features and the level of weight sparsity; Υ corr (z,y) is the Spearman correlation coefficient, which represents the hidden feature z k and performance variable y k the degree of relevance; Indicates that the predicted target is x k The i-th input variable x in the network i The connection weight of .
[0152] For the performance constraint prediction network, since a single prediction network can only obtain the potential dependent variable set of a single target variable, in order to obtain the overall performance causal skeleton of all process variables, the global loss function is minimized. To train n prediction networks simultaneously, The expression is:
[0153]
[0154] In step S5, based on the trained performance constraint prediction network, the performance causal relationship between process variables within the process is analyzed, and a process performance causal graph is generated, which is then uploaded to the cloud platform.
[0155] In this step, a process performance cause-and-effect diagram is generated, which specifically includes:
[0156] Step S51: extracting a causal relationship skeleton using a performance constraint prediction network to obtain a performance causal undirected graph.
[0157] Predicting the weights of the input layer in the network The size of x can be represented by i For the predicted target x kThe difference in input weights will be more obvious after the sparsification effect, and the variable with large input weight is more likely to have a causal relationship with the target variable. In the n prediction networks trained, we can always find a pair of variables {x u ,x v}Connection weights when they are input and target and In the causal skeleton search phase, the weights are divided As a quantitative indicator of whether there is a causal relationship between a pair of variables, the higher the weight score, the greater the possibility of a causal relationship between the corresponding variables.
[0158] Considering the input weight in the penalty term cannot completely eliminate redundant connections. In order to obtain a refined performance causal skeleton, a score threshold t is set. s Prune potential causal connections. Figure 6 As shown, in this embodiment, all the value scores are arranged in ascending order from small to large, the difference g between adjacent scores is calculated, and the maximum difference g is taken. max The score value on the left is used as the threshold t s Variable pairs with weight scores exceeding the threshold are retained as causal connections, otherwise the potential causal connections are eliminated. Based on this, the adjacency matrix A of the performance causal undirected graph can be obtained:
[0159]
[0160]
[0161] In formula (27), a u,v Represents the variable {x u ,x v} connection relationship, a u,v =1 means x u and x v There is a potential causal relationship between them, otherwise x u and x v There is no potential causal relationship between them; the adjacency matrix A contains the potential causal relationship between all variables.
[0162] Step S52: Use the complexity of the prediction model to determine the causal direction and obtain a performance causal directed graph.
[0163] In this step, in the evolution and reasoning of objective facts, the cause event (dependent variable) can affect the development and change of the effect event (effect variable). Therefore, it is believed that in the prediction network with performance constraints, the effect variable can be easily predicted by the dependent variable, that is, the model that predicts the effect variable by the dependent variable has a lower model complexity. n prediction networks with the same structure and hyperparameters were trained, and a performance causal undirected graph was obtained from them. For two variables {xu ,x v}, calculate x respectively u Predict x v Model complexity EMC u,v , and x v Predict x u Model complexity EMC v,u , including EMC u,v The calculation expression is:
[0164]
[0165] In formula (28), is the connection weight from the input layer to the first hidden layer, is the connection weight between the first hidden layer and the second hidden layer, W z = is the connection weight between the second hidden layer and the output layer, n h 、n z are the dimensions of the first hidden layer and the second hidden layer network respectively. If EMC u,v <EMC v,u , then x u is x v The dependent variable, otherwise x v is x u By traversing all undirected connected variable pairs in the causal undirected graph, we can determine the causal direction between all variables and obtain a performance causal directed graph.
[0166] Step S53: prune the performance causal directed graph to obtain a final performance causal graph.
[0167] In this step, after learning the above performance causal directed graph, the mechanism and process knowledge are further applied to prune and refine the learned causal relationships, prune out the non-directly related causal relationships, and obtain a refined performance causal directed graph to improve the logic and accuracy of the causal directed graph.
[0168] In step S6, on the end-side platform, real-time process data of the current sample point of each process is collected and preprocessed; a real-time data set for each process is constructed based on performance-related process variables, and the real-time data set is then uploaded to the corresponding edge-side platform by process.
[0169] In this step, each monitoring moment corresponds to a sampling point.
[0170] In step S7, on the edge platform, based on the trained QGRU-SAE model, the performance features in the real-time data set are extracted by process, and a process performance monitoring model is constructed, which is then uploaded to the cloud platform.
[0171] In this step, a process performance monitoring model is constructed, which specifically includes: extracting performance-related features of each process based on the trained QGRU-SAE model, and constructing T in the feature space and residual space respectively using the real-time data uploaded from the end side. 2 and SPE monitoring statistics for real-time monitoring of product performance and performance monitoring of each process.
[0172] Step S8: On the cloud-side platform, the process performance cause-effect graphs are merged to obtain a global performance cause-effect graph.
[0173] In this step, the global performance causal graph records the causal relationships among all performance-related variables, and based on the global performance causal graph, the performance causal discovery of the strip hot rolling process is completed.
[0174] Among them, in the performance causal relationship discovery on the cloud-side platform, the performance causal graphs of each process are integrated and connected using the strip hot rolling process mechanism, expert knowledge and multi-level system information stored on the cloud side, and the structural learning method based on the Bayesian Information Criterion (BIC) score is used to search for the optimal global performance causal graph.
[0175] Specifically, for the subgraph G of process p p , the BIC score is defined as follows:
[0176]
[0177] In formulas (29) and (30), D is the process data set, is the maximum likelihood parameter, is the likelihood function, Dim[·] represents the dimension, n p is the number of nodes in the subgraph, MI(·) is the mutual information, H(·) is the edge entropy, det(·) represents the discriminant, and Σ(·) is the covariance. By maximizing S BIC (G p :D) The scores are used to obtain a causal directed graph of the performance of the entire process, which is used to trace the root causes of performance-related failures in S10.
[0178] Step S9: On the cloud-side platform, the performance monitoring information of each process is obtained based on the process performance monitoring model and integrated to obtain the global performance monitoring information; based on the global performance monitoring information, it is determined whether there is a performance anomaly; if there is no anomaly, return to step S6; if there is an anomaly, go to step S10.
[0179] In this step, the real-time data set is input into the process performance monitoring model to obtain the performance monitoring information of each process at the current sampling point. The performance monitoring information includes monitoring statistics. When fusing the performance monitoring information of each process, Bayesian reasoning is used for fusion.
[0180] The threshold of the monitoring statistic is determined by the Kernel Density Estimation (KDE) method. Specifically, for process p, its process variable is The performance variables are The performance-related features extracted by the QGRU-SAE performance monitoring model are: And build monitoring statistics and SPE p Among them, n p 、m p and H p They respectively represent the process variable dimension, performance variable dimension and performance characteristic dimension under this process.
[0181] Determine whether there are performance anomalies based on global performance monitoring information, including:
[0182] The posterior probability of a performance failure occurring in process p is and the probability of performance failure throughout the entire process The calculation is as follows:
[0183]
[0184] In formulas (31) and (32), The prior probability and conditional probability of normal and abnormal product performance are and are the performance monitoring statistics and control limits of process p, N p is the number of processes in the entire production line, and α is the given confidence level. This indicates a performance-related failure on the production line. Leveraging the cloud platform's powerful computing power, we continuously optimize model parameters based on real-time data uploaded from the client side, enabling dynamic, real-time performance monitoring.
[0185] In step S10, a candidate set of fault-related variables is constructed based on the generalized reconstruction-based contribution (GRBC) graph to isolate the fault variables; and the fault propagation path is identified and the root cause variable of the fault is located based on the global performance causal graph.
[0186] In this step, the process of constructing a candidate set of performance fault-related variables includes:
[0187] Generalized contribution score GRBC of the i-th variable to performance failure in the t-th sampling point t,i for:
[0188]
[0189]
[0190] In formulas (33) and (34), i is the fault direction matrix, M is the characteristic matrix of the performance feature space, is the generalized inverse of the matrix. is the fault contribution ratio of variable i in all variables, if The i-th variable can be included in the candidate set of fault variables to isolate fault-related variables. The fault propagation path can be identified by combining the performance causal graph to locate the root cause variable of the fault.
[0191] It can be seen that the embodiment of the present invention is based on the combination of the terminal side, edge side and cloud side platforms, such as Figure 7 As shown in the figure, after the cloud-side platform detects abnormal production line performance, the platform will automatically record the time of the abnormality and the exceeding limit data of relevant performance indicators, and transmit the abnormal information to the edge platform in real time to realize the rapid positioning of the process with potential performance abnormality; after the edge platform locates the abnormal process, the platform will perform fault path analysis and root cause tracing based on the local performance cause-and-effect graph of the process, so as to accurately locate the root cause of the fault and provide reliable fault diagnosis and troubleshooting guidance for on-site operators; through the construction and deployment of the above-mentioned platform, dynamic real-time performance monitoring and fault diagnosis of the strip hot rolling process under the collaboration of cloud, edge and end are realized. The fault monitoring and root cause positioning results based on the platform feedback provide operators with a reliable basis for troubleshooting, shorten the maintenance time of system faults, and reduce the economic loss caused by product performance fluctuations.
[0192] Based on the same idea, an embodiment of the present invention also provides a performance causal discovery and fault diagnosis device for a strip hot rolling process. It should be noted that since the performance causal discovery and fault diagnosis device for a strip hot rolling process provided in this embodiment corresponds to the specific implementation method of the above-mentioned performance causal discovery and fault diagnosis method for a strip hot rolling process, the system can achieve the purpose of the present invention by executing the process steps in the specific implementation method of the above-mentioned method. Therefore, the explanations in the specific implementation method of the above-mentioned cloud-edge-end collaborative performance causal discovery and fault diagnosis method for a strip hot rolling process are also applicable to the performance causal discovery and fault diagnosis device for a strip hot rolling process. Therefore, the specific implementation method of the device will not be repeated in the following specific implementation methods of the present invention.
[0193] like Figure 7As shown, the performance causal discovery and fault diagnosis device for the hot rolling process of strip steel comprises: a data acquisition module 101, a variable screening module 102, a QGRU-SAE model construction module 103, a prediction network construction module 104, a performance causal skeleton search module 105, a performance causal discovery module 106, a process performance monitoring module 107, a global performance monitoring module 108 and a fault tracing module 109; wherein,
[0194] The data acquisition module 101 is used to acquire historical process data of each process of the strip hot rolling production line on the terminal side platform and perform preprocessing; it is also used to collect real-time process data of the current sample point of each process and perform preprocessing;
[0195] The variable screening module 102 is used to screen performance-related process variables from pre-processed historical process data, and construct a historical multidimensional dataset for each process step based on the performance-related process variables, and then upload the screened performance-related process variables and historical multidimensional dataset to the corresponding side platform by process step. It is also used to construct a real-time dataset for each process step based on the performance-related process variables, and then upload the real-time dataset to the corresponding side platform by process step.
[0196] The QGRU-SAE model construction module 103 is used to construct a performance-driven gated recurrent stacked autoencoder QGRU-SAE model based on the performance-related process variables on the side platform; and use the historical multidimensional dataset of the current process to perform model training under performance supervision;
[0197] The prediction network construction module 104 is used to obtain the historical performance characteristics of each process based on the trained QGRU-SAE model on the side platform, and to construct a performance constraint prediction network for each process. The loss function is set based on the prediction error of performance-related variables, the prediction error of performance variables, the performance correlation, and the degree of input connection sparsity, and the training is performed.
[0198] The performance causal skeleton search module 105 is used to analyze the performance causal relationship between process variables within the process based on the trained performance constraint prediction network, and generate a process performance causal graph, which is then uploaded to the cloud platform;
[0199] The performance causal discovery module 106 is used to merge the process performance causal graphs on the cloud platform to obtain a global performance causal graph;
[0200] The process performance monitoring module 107 is used to extract performance features from the real-time data set by process based on the trained QGRU-SAE model on the edge platform, and build a process performance monitoring model, which is then uploaded to the cloud platform.
[0201] The global performance monitoring module 108 is used to obtain and integrate the performance monitoring information of each process on the cloud-side platform based on the process performance monitoring model to obtain global performance monitoring information; based on the global performance monitoring information, determine whether there is a performance anomaly; if there is no anomaly, start the data acquisition module 101; if there is an anomaly, start the fault tracing module 109;
[0202] The fault tracing module 109 is used to construct a candidate set of fault-related variables based on the generalized reconstruction contribution (GRBC) graph to isolate fault variables; and to identify fault propagation paths and locate fault root variables based on the global performance causal graph.
[0203] Based on the performance causal discovery and fault diagnosis method and device of the strip hot rolling process, the embodiment of the present invention also provides a performance causal discovery and fault diagnosis system of the strip hot rolling process. Figure 8 As shown, the system includes an end-side platform, an edge-side platform and a cloud-side platform; wherein, the end-side platform corresponds to the specific equipment of each process in the hot rolling production line, the edge-side platform corresponds to each process, and the cloud-side platform corresponds to the overall hot rolling production line.
[0204] The end-side platform is used for data acquisition, preprocessing, and screening. Specifically, the end-side platform is used to obtain and preprocess the historical process data of each process in the hot-rolling production line; screen performance-related process variables from the preprocessed historical process data, and construct a historical multidimensional dataset for each process based on the performance-related process variables. The screened performance-related process variables and historical multidimensional datasets are then uploaded to the corresponding side-side platform by process. The end-side platform is also used to collect and preprocess the real-time process data of the current sample point of each process; and construct a real-time dataset for each process based on the performance-related process variables. The real-time dataset is then uploaded to the corresponding side-side platform by process.
[0205] The end-side platform here can be implemented in the form of a data management platform. In view of the non-ideal characteristics of multi-source heterogeneous data and incomplete, unbalanced and partially missing data in the hot rolling process, relying on the existing information system and data acquisition equipment, we can obtain complete data of each stage and each process of the whole life cycle of hot-rolled products, and provide a reliable data foundation for the construction of cloud-edge collaborative monitoring and diagnosis systems. The data management platform is designed to support comprehensive data management and user management. It can query and export, add and import, edit and enter, and expand and modify data item parameters for different types of databases. It can unify the management and docking of various existing data, realize the sharing and exchange of data resources, and provide a collaborative information environment and data support for subsequent performance monitoring and fault diagnosis.
[0206] The side platform is used to build a performance-driven gated recurrent stacked autoencoder (QGRU-SAE) model based on the performance-related process variables; and use the historical multidimensional data set under the current process to train the model under performance supervision; then based on the trained QGRU-SAE model, the historical performance characteristics of each process are obtained, and a performance constraint prediction network for each process is constructed, and a loss function is set based on the prediction error of the performance-related variables, the prediction error of the performance variables, the performance correlation and the degree of input connection sparsity, and training is performed; based on the trained performance constraint prediction network, the performance causal relationship between the process variables within the process is analyzed, and a process performance causal graph is generated, which is then uploaded to the cloud side platform; and based on the trained QGRU-SAE model, the performance characteristics in the real-time data set are extracted for each process, and a process performance monitoring model is constructed, which is then uploaded to the cloud side platform.
[0207] The cloud-side platform is used to merge the process performance cause-and-effect graphs to obtain a global performance cause-and-effect graph; obtain and fuse the performance monitoring information of each process based on the process performance monitoring model to obtain global performance monitoring information; determine whether there is a performance anomaly based on the global performance monitoring information; if there is no anomaly, proceed to the next sampling point; if there is an anomaly, construct a fault-related variable candidate set based on the generalized reconstruction contribution (GRBC) graph to achieve fault variable isolation; and identify the fault propagation path and locate the fault root variable based on the global performance cause-and-effect graph.
[0208] In summary, the cloud-edge-end collaborative strip hot rolling performance monitoring and fault diagnosis method and system provided in this embodiment have at least the following beneficial effects:
[0209] (1) Design a data management platform to realize the sharing and exchange of data resources on the end side and the cloud side, collect multi-dimensional real-time data of the hot rolling production line on the end side, and complete the integration and storage of data throughout the product life cycle on the cloud side, which can effectively eliminate the information omissions and incompleteness caused by the independent collection of data from each process in the traditional mode.
[0210] (2) Constructing a performance-driven gated recurrent stacked autoencoder model can fully utilize the performance information of the hot rolling production process, explore the deep-level performance characterization of the production line, and realize the construction of product performance causal graphs and dynamic performance monitoring of the hot rolling process by process and the entire process, which greatly improves the accuracy of product performance monitoring and can reliably ensure the safe and stable operation of the production line.
[0211] (3) Based on the cloud-edge-end collaborative architecture, by integrating the data resources and model algorithms of the end side, edge side and cloud side, the update and collaboration of the models between platforms in the hot rolling production process are realized. After the cloud side platform detects the production line performance abnormality, the platform will automatically record the time of the abnormality and the over-limit data of the relevant performance indicators, and transmit the abnormal information to the edge side platform in real time to realize the rapid positioning of the process with potential performance abnormality; after the edge side platform locates the abnormal process, the platform will perform fault path analysis and root cause tracing based on the local performance causal map of the process, so as to accurately locate the root cause of the fault and provide reliable fault diagnosis and troubleshooting guidance for on-site operators; through the construction and deployment of the above platform, dynamic real-time performance monitoring and fault diagnosis of the strip hot rolling process under cloud-edge-end collaboration are realized. Based on the fault monitoring and root cause positioning results fed back by the platform, a reliable fault troubleshooting basis is provided for operators, which shortens the maintenance time of system faults and reduces the economic benefit loss caused by product performance fluctuations.
[0212] Furthermore, the method disclosed in accordance with an embodiment of the present invention can also be implemented as a computer program executed by a processor, which can be stored in a computer-readable storage medium. The storage medium stores at least one instruction, which is loaded and executed by the processor to implement the method of the first embodiment described above. The computer-readable storage medium can be a ROM, random access memory, CD-ROM, magnetic tape, floppy disk, or optical data storage device, among others. The instructions stored therein can be loaded by the processor in the terminal to execute the method described above.
[0213] Furthermore, it should be noted that the present invention may be provided as a method, apparatus, or computer program product. Therefore, embodiments of the present invention may take the form of a fully or partially hardware embodiment, a fully or partially software embodiment, or an embodiment combining software and hardware aspects. Furthermore, when implemented using software, embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The computer program product comprises one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the processes or functions described in accordance with the embodiments of the present invention are fully or partially generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired connection (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium accessible by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a DVD), or a semiconductor medium. The semiconductor medium may be a solid state drive.
[0214] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the process in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0215] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for implementing the process in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0216] It should also be noted that, in this document, relational terms such as first and second are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any actual relationship or order between these entities or operations. The terms "include," "comprises," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or terminal device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. Without further limitation, an element defined by the phrase "comprising a..." does not preclude the presence of other identical elements in the process, method, article, or terminal device comprising the element. In addition, the term "and / or" is merely a description of an associative relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: the presence of A alone, the presence of A and B simultaneously, or the presence of B alone, where A and B can be singular or plural.
[0217] If the method is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0218] Finally, it should be noted that the above is a preferred embodiment of the present invention. It should be noted that although the preferred embodiment of the present invention has been described, it is clear that those skilled in the art, once they understand the basic inventive concept of the present invention, can make various improvements and modifications without departing from the principles of the present invention. Such improvements and modifications should also be considered as within the scope of protection of the present invention. Therefore, the appended claims are intended to be interpreted as including the preferred embodiment and all changes and modifications that fall within the scope of the embodiments of the present invention.
Claims
1. A method for performance causal discovery and fault diagnosis in a strip hot rolling process, characterized in that: The method comprises the following steps: Step S1: On the terminal platform, historical process data of each process of the strip hot rolling production line is obtained and pre-processed; Step S2: Filter performance-related process variables from the pre-processed historical process data, and construct a historical multidimensional dataset for each process based on the performance-related process variables. The filtered performance-related process variables and historical multidimensional datasets are then uploaded to the corresponding side platform by process. Step S3: On the side platform, based on the performance-related process variables, construct a performance-driven gated recurrent stacked autoencoder (QGRU-SAE) model; And use the historical multidimensional dataset of the current process to train the model under performance supervision; Step S4: On the side platform, based on the trained QGRU-SAE model, the historical performance characteristics of each process are obtained, and a performance constraint prediction network for each process is constructed. A loss function is set based on the prediction error of performance-related variables, the prediction error of performance variables, performance correlation, and the degree of input connection sparsity, and training is performed. Step S5: Based on the trained performance constraint prediction network, the performance causal relationship between process variables within the process is analyzed, and a process performance causal graph is generated, which is then uploaded to the cloud platform. Step S6: On the device-side platform, real-time process data from the current sample point of each process is collected and preprocessed. A real-time dataset for each process is constructed based on performance-related process variables. The real-time dataset is then uploaded to the corresponding edge-side platform by process. Step S7: On the edge platform, based on the trained QGRU-SAE model, extract performance features from the real-time data set by process, build a process performance monitoring model, and then upload it to the cloud platform; Step S8: On the cloud-side platform, the process performance cause-effect graphs are merged to obtain a global performance cause-effect graph; Step S9: On the cloud-side platform, based on the process performance monitoring model, the performance monitoring information of each process is obtained and integrated to obtain the global performance monitoring information; Determine whether there are performance anomalies based on global performance monitoring information; If there is no abnormality, return to step S6; if there is an abnormality, go to step S10; In step S10, a candidate set of fault-related variables is constructed based on the generalized reconstruction contribution graph to isolate the fault variables; and the fault propagation path is identified and the root cause variable of the fault is located based on the global performance causal graph.
2. The method according to claim 1, characterized in that The end-side platform corresponds to each device in each process of strip hot rolling; the side platform corresponds to the process; and the cloud-side platform corresponds to the global hot rolling production line.
3. The method according to claim 1, characterized in that The preprocessing in step S1 includes: Step S11, normalizing the historical process data to eliminate the dimensional differences of multi-source data; Step S12: using a multiple interpolation algorithm to fill in the missing data in the normalized historical process data.
4. The method according to claim 1, wherein In step S2, the process variables related to performance are screened, specifically including: Step S21, calculating the Pearson correlation coefficient of the variables based on the knowledge of the hot rolling process mechanism and the collected data; Step S22 : setting a coefficient threshold, and taking variables corresponding to Pearson coefficients greater than the coefficient threshold as screened performance-related process variables.
5. The method according to claim 4, characterized in that The coefficient threshold is set to the empirical threshold ρ th =0.
4.
6. The method according to claim 1, characterized in that In step S3, a performance-driven gated recurrent stacked autoencoder (QGRU-SAE) model is constructed and trained, specifically including: Assume that the process variable is X=[x1,x2,…,x n ]∈R N×n , the performance variable is Y=[y1,y2,…,y m ]∈R N×m , based on the gated recurrent unit GRU, a performance-supervised QGRU is constructed, and the performance information is used to guide the feature learning of the forget gate and hidden state. The calculation process of the performance-supervised QGRU is as follows: r(t)=f r (W r *[h(t-1),X(t)]+U r *Y(t)+b r ) (15) z(t)=f z (W z *[h(t-1),X(t)]+U z *Y(t)+b z ) (16) In formulas (15)-(18), r(t) and z(t) are the values of the reset gate and the update gate at time t, respectively; h(t) are the candidate state vector and output vector of the QGRU unit at time t; f(·) represents the activation function; tanh represents the hyperbolic tangent activation function when updating the candidate state vector; X(t) represents [x1, x2, …, x n ]The value of the variable at time t, Y(t) represents [y1,y2,…,y m ]The value of the variable at time t, X(t) and Y(t) serve as the input of the QGRU unit at time t; {W r ,b r }、{W z ,b z }、{W h ,b h } are the network parameters of reset gate, update gate and hidden gate respectively; U r 、U z 、U h are the training weights of the performance variables in the reset gate, update gate, and hidden gate respectively; ⊙ represents the element-wise product; The QGRU-SAE n neural network composed of n layers of QGRU stacked layer by layer is pre-trained layer by layer to extract deep data features. First, the process variable X and performance variable Y are input into QGRU-SAE 1. In the encoding part, the first layer hidden layer feature h is obtained through QGRU learning. 1 , in the decoding part, use h 1 To reconstruct the process variable X and performance variable Y, we get the reconstructed And the first reconstruction By minimizing Implement pre-training of QGRU-SAE 1; the first reconstruction expression and error are: In formulas (19)-(20), f(·) represents the activation function of the decoding layer, Network parameters corresponding to the process variable and performance variable decoding layers respectively; is the training parameter set of QGRU-SAE 1; The hidden layer features h of QGRU-SAE 1 1 And the performance variable Y is input into QGRU-SAE 2, and the deep hidden layer feature h is obtained by QGRU learning at the encoding layer 2 , in the decoding part, use h 2 To reconstruct the hidden layer features h 1 and performance variable Y, and obtain the reconstructed And the second reconstruction By minimizing Implement pre-training of QGRU-SAE 2; the second reconstruction expression and error are: In formulas (21) and (22), f(·) represents the activation function of the decoding layer, The network parameters corresponding to the hidden layer features and the performance variable decoding layer respectively; is the training parameter set of QGRU-SAE 2; Similarly, the stacked QGRU-SAE n network is pre-trained layer by layer under the supervision of the performance variable Y.
7. The method according to claim 1, characterized in that In step S4, a loss function is set based on the prediction error of performance-related variables, the prediction error of performance variables, performance correlation, and the degree of input connection sparsity. The loss function expression of a single prediction network is as follows: In formula (25), represents the loss function of the k-th prediction network; Corresponding to the predicted value and true value of the target process variable and performance variable respectively; α and β are the hyperparameters corresponding to the penalty term, which are used to control the performance correlation degree of network features and the level of weight sparsity; Υ corr (z,y) is the Spearman correlation coefficient, which represents the hidden feature z k and performance variable y k the degree of relevance; Indicates that the predicted target is x k The i-th input variable x in the network i The connection weight of For the performance constraint prediction network, since a single prediction network can only obtain the potential dependent variable set of a single target variable, in order to obtain the overall performance causal skeleton of all process variables, the global loss function is minimized. To train n prediction networks simultaneously, The expression is:
8. The method according to claim 1, characterized in that In step S5, a process performance cause-and-effect diagram is generated, which specifically includes: Step S51, using the performance constraint prediction network to extract the causal relationship skeleton to obtain a performance causal undirected graph; Step S52: using the complexity of the prediction model to determine the causal direction and obtain a performance causal directed graph; Step S53: prune the performance causal directed graph to obtain a final performance causal graph.
9. A device for performance causal discovery and fault diagnosis in a hot rolling process of a steel strip, characterized in that: The device includes: a data acquisition module 101, a variable screening module 102, a QGRU-SAE model construction module 103, a prediction network construction module 104, a performance causal skeleton search module 105, a performance causal discovery module 106, a process performance monitoring module 107, a global performance monitoring module 108 and a fault tracing module 109; wherein, The data acquisition module 101 is used to acquire historical process data of each process of the strip hot rolling production line on the terminal side platform and perform preprocessing; it is also used to collect real-time process data of the current sample point of each process and perform preprocessing; The variable screening module 102 is used to screen performance-related process variables from pre-processed historical process data, and construct a historical multidimensional dataset for each process step based on the performance-related process variables, and then upload the screened performance-related process variables and historical multidimensional dataset to the corresponding side platform by process step. It is also used to construct a real-time dataset for each process step based on the performance-related process variables, and then upload the real-time dataset to the corresponding side platform by process step. The QGRU-SAE model construction module 103 is used to construct a performance-driven gated recurrent stacked autoencoder QGRU-SAE model based on the performance-related process variables on the side platform; and use the historical multidimensional dataset of the current process to perform model training under performance supervision; The prediction network construction module 104 is used to obtain the historical performance characteristics of each process based on the trained QGRU-SAE model on the side platform, and to construct a performance constraint prediction network for each process. The loss function is set based on the prediction error of performance-related variables, the prediction error of performance variables, the performance correlation, and the degree of input connection sparsity, and the training is performed. The performance causal skeleton search module 105 is used to analyze the performance causal relationship between process variables within the process based on the trained performance constraint prediction network, and generate a process performance causal graph, which is then uploaded to the cloud platform; The performance causal discovery module 106 is used to merge the process performance causal graphs on the cloud platform to obtain a global performance causal graph; The process performance monitoring module 107 is used to extract performance features from the real-time data set by process based on the trained QGRU-SAE model on the edge platform, and build a process performance monitoring model, which is then uploaded to the cloud platform. The global performance monitoring module 108 is used to obtain and integrate the performance monitoring information of each process on the cloud-side platform based on the process performance monitoring model to obtain global performance monitoring information; based on the global performance monitoring information, determine whether there is a performance anomaly; if there is no anomaly, start the data acquisition module 101; if there is an anomaly, start the fault tracing module 109; The fault tracing module 109 is used to construct a candidate set of fault-related variables based on the generalized reconstruction contribution (GRBC) graph to isolate fault variables; and to identify fault propagation paths and locate fault root variables based on the global performance causal graph.
10. A performance causal discovery and fault diagnosis system for a hot rolling process of a strip steel, characterized in that: The system includes an end-side platform, an edge-side platform, and a cloud-side platform; wherein the end-side platform corresponds to the specific equipment of each process in the hot rolling production line, the edge-side platform corresponds to each process, and the cloud-side platform corresponds to the entire hot rolling production line; The end-side platform is used to obtain historical process data of each process of the strip hot rolling production line and perform preprocessing; screen performance-related process variables from the preprocessed historical process data, and build a historical multidimensional data set for each process based on the performance-related process variables, and then upload the screened performance-related process variables and historical multidimensional data sets to the corresponding side-side platform by process; it is also used to collect real-time process data of the current sample point of each process and perform preprocessing; and build a real-time data set for each process based on the performance-related process variables, and then upload the real-time data set to the corresponding side-side platform by process; The side platform is used to build a performance-driven gated recurrent stacked autoencoder QGRU-SAE model based on the performance-related process variables; and use the historical multidimensional data set under the current process to train the model under performance supervision; then based on the trained QGRU-SAE model, the historical performance characteristics of each process are obtained, and a performance constraint prediction network for each process is constructed, and a loss function is set based on the prediction error of the performance-related variables, the prediction error of the performance variables, the performance correlation and the degree of input connection sparsity, and training is performed; based on the trained performance constraint prediction network, the performance causal relationship between the process variables in the process is analyzed, and a process performance causal graph is generated, and then uploaded to the cloud-side platform; and based on the trained QGRU-SAE model, the performance characteristics in the real-time data set are extracted for each process, and a process performance monitoring model is constructed, and then uploaded to the cloud-side platform; The cloud-side platform is used to merge the process performance cause-and-effect graphs to obtain a global performance cause-and-effect graph; obtain and fuse the performance monitoring information of each process based on the process performance monitoring model to obtain global performance monitoring information; determine whether there is a performance anomaly based on the global performance monitoring information; if there is no anomaly, proceed to the next sampling point; if there is an anomaly, construct a fault-related variable candidate set based on the generalized reconstruction contribution (GRBC) graph to achieve fault variable isolation; and identify the fault propagation path and locate the fault root variable based on the global performance cause-and-effect graph.
Citation Information
Cited By
Industrial process fault detection method based on space-time causal graph auto-encoder
CN120974245A