A Distributed Canonical Correlation Analysis Process Monitoring Method and Device

By blocking and modeling large-scale industrial processes, using topological matrix and residual generators, the problems of low monitoring sensitivity and inaccurate fault positioning in the prior art are solved, and more efficient fault detection and positioning are achieved.

CN116048024BActive Publication Date: 2025-07-04EAST CHINA UNIV OF SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310059039.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-18
Publication Date
2025-07-04
Estimated Expiration
2043-01-18

AI Technical Summary

Technical Problem

The existing distributed process monitoring methods have shortcomings in monitoring sensitivity and fault location, making it difficult to accurately detect faults in large-scale industrial processes.

Method used

A distributed typical correlation analysis method based on partial sub-block communication is adopted. By blocking the training sample data, building a topological matrix, typical correlation analysis modeling, a residual generator is generated, and statistics are calculated to achieve fault detection.

Benefits of technology

It improves monitoring sensitivity, reduces computational complexity and data transmission load, and can accurately detect faults and initially locate fault ranges.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116048024B_ABST
    Figure CN116048024B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of distributed process detection, and more specifically, to a distributed canonical correlation analysis process monitoring method and device based on partial sub-block communication. The present invention includes: dividing training sample data into blocks and determining a topological matrix; performing canonical correlation analysis modeling on each sub-block according to the topological matrix to construct a residual generator; generating a residual vector according to the residual generator and calculating a statistic; performing corresponding partitioning on new sampled data; substituting the partitioned sampled data into the corresponding canonical correlation analysis model respectively to obtain corresponding canonical correlation components; performing fault detection on each sub-block and obtaining the operating state of the sub-block; calculating an overall statistical index FD of all sub-blocks, and determining whether a fault occurs in the process by comparing the overall statistical index FD with a control limit. The present invention has the advantages of improving the sensitivity of monitoring, reducing the computational complexity of process monitoring and the load of data transmission, and realizing the preliminary range positioning of faults, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of distributed process detection, and more specifically, to a distributed canonical correlation analysis process monitoring method and device based on partial sub-block communication. Background Art

[0002] In recent years, with the development of modern industrial processes, large-scale processes with multiple interconnected operating units have become very popular, and monitoring such processes has increasingly become the focus of attention.

[0003] Large-scale plant-wide processes such as chemical plants and oil refineries are basically composed of several different operating units and workshops, and even the production operations are jointly carried out by several factories in different regions. Although the production scale is very large, in order to ensure a safe operating environment and produce high-quality products, it is very necessary to monitor the entire production process and even the entire factory.

[0004] With the development of the scale of modern industrial systems, the process data collected has the characteristics of being large in quantity and high in dimension. Extracting information that can reflect the essence of process production from numerous high-dimensional measurement data is an important factor in process monitoring.

[0005] In this regard, on the one hand, researchers have adopted various data dimensionality reduction methods. By reducing the dimension of the data, the high-dimensional data space is mapped to a low-dimensional feature space, while preserving the vast majority of the information in the original data and removing the redundant information between the data. The most common and most used ones are the principal component analysis method and the independent component analysis method.

[0006] Principal Component Analysis (PCA) is a linear dimensionality reduction method first proposed by Hotelling in 1933. The PCA method orthogonally transforms the originally correlated variables to obtain linearly uncorrelated principal components. While the transformed principal components contain as much information of the original data as possible, it is hoped that the correlation between the principal components is as small as possible. Therefore, the principal components extracted by PCA can not only denoise the signal but also reduce the dimensionality of the data with a small amount of information loss, which is beneficial to process monitoring. However, the PCA method is applicable to data that follows a Gaussian distribution. For cases where the data does not satisfy the Gaussian distribution, the Independent Component Analysis (ICA) method is often used. Independent Component Analysis was first proposed by Jutten and Herault and is used to recover the basic independent source signals from the observed linearly mixed signals. The goal of ICA is to decompose the multivariate data into a linear combination of statistically independent components. ICA assumes that the source information follows a non-Gaussian distribution and is mutually independent, so it is applicable to the monitoring of non-Gaussian processes. In addition to process monitoring, Independent Component Analysis has also been widely used in many fields, such as image processing, pattern recognition, data mining, etc.

[0007] On the other hand, for complex large-scale processes, people adopt the idea of "breaking the whole into parts", that is, dividing an entire production process into several scattered units and monitoring them separately, namely distributed process monitoring. Distributed process monitoring is often applied to the process monitoring of the entire factory.

[0008] First, the entire factory process is decomposed into different blocks. Then, statistical monitoring models are constructed in different blocks. Finally, the monitoring results of different blocks are integrated through a decision fusion strategy. Currently, many applications have been made of the distributed process monitoring method based on multi-block statistical modeling. An actual industrial process consists of multiple interconnected operating units or devices, and there are connections and mutual influences between different operating units. When the production process is operating in a normal state, the relationship between the units shows a certain pattern, and the correlation between the measurement data of the two is relatively stable.

[0009] However, the existing distributed process monitoring methods based on multi-block statistical modeling only extract features and establish monitoring models based on the measurement data of the unit blocks, ignoring the continuity of the entire production process, resulting in low monitoring sensitivity and difficulty in accurate fault location. Summary of the Invention

[0010] The object of the present invention is to provide a distributed canonical correlation analysis process monitoring method based on partial sub-block communication to solve the problems of low monitoring sensitivity and difficulty in accurate fault location of the existing distributed process monitoring methods.

[0011] To achieve the above object, the present invention provides a distributed canonical correlation analysis process monitoring method, including an offline modeling stage and an online monitoring stage. The offline modeling stage specifically includes the following steps:

[0012] Step S1: Divide the training sample data into blocks and determine a topological matrix for representing the connection situation between sub-blocks;

[0013] Step S2: Based on the connection situation of the topological matrix, perform canonical correlation analysis modeling on the data of the current sub-block and the sub-blocks from the connections, extract the canonical correlation components and retain the projection matrix, and construct a residual generator according to the correlation relationship between sub-blocks;

[0014] Step S3: Generate a residual vector according to the residual generator and calculate the statistic;

[0015] The online monitoring stage specifically includes the following steps:

[0016] Step S4: Divide the new sampled data into corresponding blocks according to the method in the offline training stage;

[0017] Step S5: Substitute the sampled data obtained after being divided in Step S4 into the corresponding canonical correlation analysis models respectively, and obtain the canonical correlation components of the new sample data through conversion by the projection matrix;

[0018] Step S6: Perform fault detection on each sub-block and obtain the operating state of the sub-block;

[0019] Step S7: Calculate the overall statistical index FD of all sub-blocks, and determine whether there is a fault in the process by comparing the overall statistical index FD with the control limit, so as to realize online monitoring of the operating state of the process.

[0020] In one embodiment, Step S1 further includes:

[0021] Use a block division method based on process knowledge to divide the training sample data, and divide the process according to the connection situation of the device equipment and operation units in the actual industrial production process, using prior knowledge and expert experience.

[0022] In one embodiment, in Step S1, the topological matrix C is used to describe the connection relationship between sub-blocks. Among them, each element represents whether there is information interaction between the corresponding two sub-blocks. If so, it is set to 1; if not, it is 0.

[0023] In one embodiment, in Step S2, constructing the residual generator specifically includes the following steps:

[0024] Step S201: Perform canonical correlation analysis modeling on each pair of sub-blocks according to the topological matrix C. Take sub-blocks X1 and X2 as the input and output variables of the canonical correlation analysis algorithm respectively, and extract the canonical correlation components of sub-blocks X1 and X2. The expressions of the canonical correlation components of sub-blocks X1 and X2 are respectively:

[0025] u 12 = A 12 T X1

[0026] v 12 = B 12 T X2

[0027] Where A 12 and B 12 are the projection matrices of the two sub-blocks respectively;

[0028] Step S202: Construct a residual generator based on the canonical correlation components of each pair of sub-blocks. The expression of the residual generator between sub-blocks X1 and X2 is:

[0029] r 12 = B 12 T X2 - Λ k,12 A 12 T X1

[0030] Where Λ k,12 is the diagonal matrix of the canonical correlation coefficients of the first k pairs of correlated variables of sub-blocks X1 and X2, and the value of k is obtained by cross-validation.

[0031] In one embodiment, in step S201, perform canonical correlation analysis modeling on each pair of sub-blocks according to the topological matrix C:

[0032] When a sub-block has multiple connected sub-blocks, for the i-th sub-block, construct canonical correlation analysis models;

[0033] Where c ij is an element of the topological matrix C, and B is the number of sub-blocks.

[0034] In one embodiment, the statistic in step S3 has the corresponding expression:

[0035]

[0036] Where is the statistic of sub-block X1, n is the degree of freedom, r 12is the residual vector between sub-block X1 and sub-block X2. Correspondingly, the statistic of sub-block X2 is calculated by swapping the input and output in the residual generator.

[0037] In one embodiment, step S6 further includes the following steps:

[0038] Step S601: The current sub-block receives the canonical correlation components from the connected sub-block and generates a residual vector;

[0039] Step S602: Calculate the value of the statistic according to the residual vector. When there is one connected sub-block for the same sub-block, compare the statistic with the control limit;

[0040] When there are multiple connected sub-blocks for the same sub-block, multiple statistics are obtained. The decision fusion method is used to fuse the multiple statistics of the same sub-block to obtain the monitoring statistic. Determine whether the value of the monitoring statistic of the sub-block exceeds the threshold. If so, it indicates that the current sub-block has a fault. If not, the operating state of the current sub-block is normal, and the monitoring results of canonical correlation analysis for each sub-block are obtained.

[0041] In one embodiment, in step S602, Bayesian inference is selected as the decision fusion method. The expression of the monitoring statistic of the i-th sub-block is:

[0042]

[0043] Where is the monitoring statistic of the i-th sub-block, P(c ij ) = c ij is the coefficient representing the connection relationship between sub-blocks based on the topological matrix C, x new is the online data, is the conditional probability of the abnormal operating state of the process.

[0044] In one embodiment, in step S602, the control limit corresponding to the i-th sub-block is the prior probability (1 - α) of the occurrence of abnormal working conditions of the process:

[0045] Determine whether the monitoring statistic of the sub-block exceeds the control limit corresponding to the monitoring statistic. If so, it is considered that the sub-block has a fault. If not, it is considered that the sub-block is operating under normal working conditions.

[0046] In one embodiment, the overall statistical index FD in step S7 has the corresponding expression:

[0047]

[0048] Where is the monitoring statistic of the B sub-blocks, C Lr,B is the control limit corresponding to the B-th sub-block.

[0049] To achieve the above object, the present invention provides a distributed canonical correlation analysis process monitoring device, comprising:

[0050] A memory for storing instructions executable by a processor;

[0051] A processor for executing the instructions to implement the method as described in any one of the above.

[0052] To achieve the above object, the present invention provides a computer-readable medium having computer instructions stored thereon, wherein when the computer instructions are executed by a processor, the method as described in any one of the above is executed

[0053] The distributed canonical correlation analysis process monitoring method and device based on partial sub-block communication proposed by the present invention have the advantages of improving the sensitivity of monitoring, reducing the computational complexity of process monitoring and the load of data transmission, and realizing the preliminary range positioning of faults, etc. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] The above and other features, properties, and advantages of the present invention will become more apparent from the following description in conjunction with the drawings and embodiments, in which like reference numerals always represent like features, wherein:

[0055] Figure 1 Discloses a flowchart of a distributed canonical correlation analysis process monitoring method according to an embodiment of the present invention;

[0056] Figure 2 Discloses a schematic diagram of the principle of a distributed canonical correlation analysis process monitoring method according to an embodiment of the present invention;

[0057] Figure 3 Discloses a flowchart of a No. 1 reference simulation model according to an embodiment of the present invention;

[0058] Figure 4 Discloses a schematic diagram of the fault detection result of sub-block 1 for fault 1 according to an embodiment of the present invention;

[0059] Figure 5 Discloses a schematic diagram of the fault detection result of sub-block 1 for fault 2 according to an embodiment of the present invention;

[0060] Figure 6 Discloses a schematic diagram of the fault detection result of sub-block 2 for fault 1 according to an embodiment of the present invention;

[0061] Figure 7 Discloses a schematic diagram of the fault detection result of sub-block 2 for fault 2 according to an embodiment of the present invention;

[0062] Figure 8Disclosed is a schematic diagram of the fault detection result of sub-block 3 for fault 1 according to an embodiment of the present invention;

[0063] Figure 9 Disclosed is a schematic diagram of the fault detection result of sub-block 3 for fault 2 according to an embodiment of the present invention;

[0064] Figure 10 Disclosed is a schematic diagram of the fault detection result of the whole for fault 1 according to an embodiment of the present invention;

[0065] Figure 11 Disclosed is a schematic diagram of the fault detection result of the whole for fault 2 according to an embodiment of the present invention. Detailed implementation manners

[0066] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the invention and are not used to limit the invention.

[0067] To enable those skilled in the art to understand the features and effects of the present invention, the following provides a general description and definition of the terms and phrases mentioned in the specification and claims. Unless otherwise specified, all technical and scientific terms used herein have the ordinary meanings understood by those skilled in the art for the present invention. In case of conflict, the definitions in this specification shall prevail.

[0068] In this article, for the sake of concise description, all possible combinations of all technical features in each implementation or embodiment are not described. Therefore, as long as there is no contradiction in the combination of these technical features, the technical features in each implementation or embodiment can be combined arbitrarily, and all possible combinations should be considered as the scope described in this specification.

[0069] Figure 1 Disclosed is a flowchart of a distributed canonical correlation analysis process monitoring method according to an embodiment of the present invention, Figure 2 Disclosed is a schematic diagram of the principle of a distributed canonical correlation analysis process monitoring method according to an embodiment of the present invention. As Figure 1 and Figure 2 shown, a distributed canonical correlation analysis process monitoring method based on partial sub-block communication proposed by the present invention includes an offline modeling stage and an online monitoring stage. The offline modeling stage specifically includes the following steps:

[0070] Step S1: Divide the training sample data into blocks and determine the topological matrix C used to represent the connection situation between sub-blocks;

[0071] Step S2: Perform canonical correlation analysis (CCA) modeling on the data of the current sub-block and the sub-blocks from the connections according to the connection situation of the topological matrix C, extract the canonical correlation components and retain the projection matrix, and construct a residual generator according to the correlation relationship between the sub-blocks;

[0072] Step S3: Generate a residual vector according to the residual generator and calculate the statistic T 2 ;

[0073] The online monitoring stage specifically includes the following steps:

[0074] Step S4: Perform corresponding block division on the new sampled data according to the method in the offline training stage;

[0075] Step S5: Substitute the sampled data after block division obtained in Step S4 into the corresponding canonical correlation analysis model respectively, and obtain the canonical correlation components of the new sample data through conversion by the projection matrix;

[0076] Step S6: Perform fault detection on each sub-block and obtain the operating state of the sub-block;

[0077] Step S7: Calculate the overall statistical index FD of all sub-blocks, and judge whether there is a fault in the process by comparing the overall statistical index FD with the control limit, so as to realize online monitoring of the operating state of the process.

[0078] The present invention provides a distributed canonical correlation analysis process monitoring method based on partial sub-block communication. Considering the physical connection situation of operation units in the actual industrial process, there are connections and mutual influences between different operation units, and a distributed CCA process monitoring of partial sub-block communication is proposed:

[0079] First, perform block division in combination with the process background and relevant knowledge, and propose a topological matrix for describing the internal connection and information interaction between sub-blocks. Then, according to the topological matrix, perform CCA modeling on the sub-blocks with information transfer, extract the relevant components of the sub-blocks and use them to construct a residual generator;

[0080] Then, calculate the statistic based on the residual vector to monitor the process.

[0081] On the one hand, the residual generator can avoid directly transmitting the original data, thereby reducing the communication load of the data while improving the security of data transmission; on the other hand, by means of block division, establishing a local model can better detect small faults, improve the sensitivity of monitoring, and contribute to the location of faults.

[0082] For a large-scale production process with multiple connection units, the present invention monitors the operating state of the process by exploring whether the correlation relationship between the measurement data of different units changes. Compared with the principal component analysis technique (PCA) and independent component analysis method (ICA) that only explore the information of a single data set, canonical correlation analysis (CCA) can explore the correlation relationship between two data sets and find the mapping that makes the projections of the two sets of variables most correlated.

[0083] Therefore, the present invention selects CCA to model the relationship between different sub-blocks, that is, to establish a fault detection model between sub-blocks. Taking the solution method based on Singular Value Decomposition (SVD) as an example, the specific algorithm flow of the CCA algorithm is as follows:

[0084] Input two sets of vectors x = [x1, x2,..., x m T and y = [y1, y2,..., y l T ;

[0085] (1) Calculate the variances of the two sets of vectors and their corresponding covariances respectively:

[0086] Σ xx = cov(X, X)

[0087] Σ yy = cov(Y, Y)

[0088] Σ xy = cov(X, Y)

[0089] Σ yx = cov(Y, X)

[0090] Among them, Σ xx is the autocovariance of X, Σ yy is the autocovariance of Y, Σ xy is the covariance between X and Y, Σ yx is the covariance between Y and X;

[0091] (2) Calculate the H matrix, and the expression of the H matrix is:

[0092] H = Σ yx -1 / 2 Σ xy Σ yy -1 / 2 ;

[0093] (3) Apply singular value decomposition to the H matrix to obtain a series of singular values λ and the corresponding left and right singular vectors u and v;

[0094] ​​(4) Calculate the projection vectors a and b of the two groups of variables x and y respectively. The expressions for the projection vectors a and b are:

[0095] a = ∑ yx -1 / 2 u;

[0096] b = ∑ yy -1 / 2 v.

[0097] The specific algorithm flow of process monitoring based on the CCA algorithm is as follows:

[0098] (1) Based on the goal of exploring the correlation relationship between two groups of variables, the CCA method extracts the relevant features from the two groups of variables, namely the canonical correlation components;

[0099] (2) For the two groups of relevant features, considering the measurement noise that is difficult to avoid in the process, the expression for their correlation is:

[0100] B T y = Λ k A T x + E

[0101] where A = [a1, a2,..., a k and B = [b1, b2,..., b k are both projection matrices, Λ k = diag(λ1, λ2,..., λ k ) is a diagonal matrix, and E is the part of B T y that is uncorrelated with A T x;

[0102] (3) The CCA method assumes that the two groups of input and output variables x and y are both continuously distributed according to a multivariate normal distribution, and E is also normally distributed. The expression for the residual vector is defined as:

[0103] r = B T y - Λ k A T x

[0104] where r is the residual vector.

[0105] (4) Use the residual vector to construct a statistic T 2 , to monitor the change of the process state space. The expression for the statistic T 2 is:

[0106]

[0107] where n is the degree of freedom, and the statistic T 2The control limits are obtained through kernel density estimation or obtained according to the F-distribution it follows at a specific confidence level α, and the statistic T 2 The expression for the control limits of is:

[0108]

[0109] where is the control limit, and F α (k,n - k) is the F-distribution with degrees of freedom k and n - k.

[0110] In addition, in some methods, the features extracted from the input variables and output variables will respectively construct statistics to monitor the changes in the input-output space of the system, and the expressions for their statistics are respectively:

[0111]

[0112]

[0113] where A T x and B T y both follow the standard normal distribution with zero mean and unit variance, and the acquisition of their control limits is as described above. During process detection, when the value of the statistic is greater than its corresponding control limit, it can be considered that there is a fault in the current process; otherwise, it can be considered that the process is operating normally without faults.

[0114] Next, in combination with Figure 1 and Figure 2 , the above steps of the present invention will be described in detail. It should be understood that within the scope of the present invention, the above technical features of the present invention and the technical features specifically described below (such as in the embodiments) can be combined with each other and are interrelated to form a preferred technical solution.

[0115] Step S1: Divide the training sample data into blocks and determine the topological matrix used to represent the connection situation between sub-blocks.

[0116] Use the block division method based on process knowledge to divide the training sample data, which specifically includes the following steps:

[0117] According to the connection situation of the device equipment and operation units in the actual industrial production process, use prior knowledge and expert experience to divide the process. The expert experience includes physical constraints, process topology, and feedback constraints to achieve a preliminary range positioning of the fault after the fault is detected.

[0118] There are more and more operation units and devices in the process. For the monitoring of large-scale processes, a multi-block statistical modeling method is mostly used for distributed process monitoring, and the entire factory process needs to be decomposed into several different parts.

[0119] The framework of distributed modeling is widely used in the monitoring of large-scale processes, thanks to the following advantages:

[0120] First of all, distributed modeling can improve the fault tolerance of the process. When individual blocks have problems, the remaining blocks can still be used for monitoring;

[0121] Secondly, under the distributed modeling framework, the sensitivity analysis of the process monitoring system will be significantly improved;

[0122] Moreover, if a fault is detected in the process, the fault location can be determined within a specific block or several blocks, and further diagnostic steps can be carried out to determine the root cause of the fault;

[0123] Finally, under the distributed modeling framework, modeling each sub-block separately can effectively reduce the computational complexity of the monitoring process.

[0124] The methods of dividing process data into blocks can be divided into two categories: process-knowledge-based methods and data-based methods.

[0125] The data-based method is to divide according to the characteristics of the data itself. For example, it is divided into linear and non-linear variables, or process variables are divided according to whether they are related to quality or not.

[0126] The process-knowledge-based method is to use prior knowledge and expert experience, including physical constraints, process topology, and feedback constraints, to decompose the process.

[0127] In this embodiment, for the process-knowledge-based block division method, according to the device and equipment of the actual industrial production process and the connection situation of the operation units, the variables of the entire industrial system are divided into B blocks:

[0128] X = [X1, X2, …, X B

[0129] where X is the variable of the entire industrial system, and X B is the Bth sub-block;

[0130] For the method of dividing blocks based on process knowledge, the division difficulty is different for different production processes. Such a division has a certain feasibility and is also conducive to the preliminary scope positioning of faults after detecting faults;

[0131] In this embodiment, after dividing the process into B sub-blocks according to mechanism knowledge and expert experience, the corresponding process topology matrix C is defined.

[0132] ​The topological matrix C is proposed to describe the connection relationship between sub - blocks. Considering that in actual industrial processes, different technological processes have diverse structures, not all sub - blocks can transmit information to each other, nor do all sub - blocks have internal connections such as hierarchical progression in structure.

[0133] Therefore, compared with the distributed CCA process monitoring where all sub - blocks interact pairwise, the present invention proposes a topological matrix C to describe the connection mode between sub - blocks, and based on the topological matrix C, models the partial sub - blocks for communication, thereby further reducing the communication load of data transmission, and at the same time, the established distributed CCA process monitoring is more in line with the actual industrial process.

[0134] Each element in the topological matrix C represents whether there is information interaction between the corresponding two sub - blocks. If so, it is set to 1; if not, it is 0.

[0135] Since the element represents the relationship between two sub - blocks, the elements on the diagonal are all set to 0.

[0136] Step S2: Perform canonical correlation analysis modeling on the data of the current sub - block and the sub - blocks from the connections according to the connection situation of the topological matrix, extract the canonical correlation components and retain the projection matrix, and construct a residual generator according to the correlation relationship between sub - blocks.

[0137] Considering the internal connection and correlation relationship between different sub - blocks, this part of the information is used to establish the data model to help improve the performance of process monitoring.

[0138] In this embodiment, the canonical correlation analysis (CCA) algorithm is used to explore the correlation relationship between sub - blocks and establish a fault detection model, and a residual generator is constructed according to the correlation relationship between sub - blocks, which specifically includes the following steps:

[0139] Step S201: Perform CCA modeling on each pair of sub - blocks according to the topological matrix C. The CCA method is based on the goal of exploring the correlation relationship between two sets of variables, and extracts the relevant features from the two sets of variables, that is, the canonical correlation components. For the B sub - blocks of the process, their corresponding CCA models are established according to the topological matrix C. For sub - blocks X1 and X2, as the input and output variables of the CCA algorithm, the canonical correlation components of sub - blocks X1 and X2 are extracted respectively. The expressions of the canonical correlation components of sub - blocks X1 and X2 are:

[0140] u 12 =A 12 T X1

[0141] v 12 =B 12 T X2

[0142] Among them, A 12 and B 12 are the projection matrices of two sub - blocks respectively;

[0143] When a sub - block has multiple connected sub - blocks, for the i - th sub - block, construct CCA models, where c ij is an element of the topological matrix C, and B is the number of sub - blocks.

[0144] Step S202: Construct a residual generator based on the canonical correlation components of pairwise sub - blocks. The expression of the residual generator between sub - block X1 and sub - block X2 is:

[0145] r 12 = B 12 T X2 - Λ k,12 A 12 T X1

[0146] where Λ k,12 is the diagonal matrix of the canonical correlation coefficients of the first k pairs of correlated variables between sub - block X1 and sub - block X2, and the selection of k is obtained by the method of cross - validation.

[0147] Step S3: Generate a residual vector according to the residual generator and calculate the statistic.

[0148] Construct a statistic based on the residual vector to monitor the change of the process state space. The expression of the statistic is:

[0149]

[0150] where, is the statistic of sub - block X1, n is the degree of freedom, r 12 is the residual vector between sub - block X1 and sub - block X2. Correspondingly, the statistic of sub - block X2 is calculated by swapping the input and output in the residual generator.

[0151] Similarly, for other sub - blocks, such as sub - block 1 and 3, sub - block 2 and 4, etc., the CCA algorithm can be used to extract the correlated features between two sub - blocks and used for the construction of the fault monitoring model.

[0152] As described above, perform CCA modeling on pairwise sub - blocks respectively. Corresponding to the actual industrial process, that is, there is communication and data transmission between these two sub - blocks. Compared with the existing centralized process monitoring that transmits all measurement data to a certain center, it can reduce the transmission distance and communication volume when data is transmitted between nodes, and construct a statistic based on the residual vector for process state monitoring. Only the corresponding k canonical correlation components need to be transmitted between sub - blocks.

[0153] Therefore, on the one hand, compared with all the original data of the transmission node, the communication load is reduced; on the other hand, transmitting non-original data can, to a certain extent, enhance the data security. Since the industrial process has a wide range and a complex structure, some measurement data can only be stored and used locally and are not suitable for sharing.

[0154] The calculation methods corresponding to steps S4 to S5 in the online monitoring stage are the same as those in the offline modeling stage, which will not be elaborated here.

[0155] In step S5, for the i-th sub-block, when new sampling data arrives, it is respectively substituted into the canonical correlation analysis model to obtain detection results.

[0156] Step S6: Perform fault detection on each sub-block and obtain the operating state of the sub-block.

[0157] Obtain the operating states of each sub-block, which specifically includes the following steps:

[0158] Step S601: The current sub-block receives the canonical correlation components from the connected sub-blocks and generates a residual vector;

[0159] Step S602: Calculate the value of the statistic according to the residual vector. When there is one connected sub-block for the same sub-block, compare the statistic with the control limit;

[0160] When there are multiple connected sub-blocks for the same sub-block, multiple statistics are obtained, that is, the detection results of the CCA model. The decision fusion method is used to fuse the multiple statistics of the same sub-block to obtain the monitoring statistic, and determine whether the value of the monitoring statistic of the sub-block exceeds the threshold. If so, it indicates that a fault occurs in the current sub-block; if not, the operating state of the current sub-block is normal, and the CCA monitoring results of each sub-block are obtained.

[0161] In this embodiment, Bayesian inference is selected as the decision fusion method to fuse the multiple statistics of the same sub-block to obtain the monitoring statistic. The expression of the monitoring statistic of the i-th sub-block is:

[0162]

[0163] where is the monitoring statistic of the i-th sub-block, P(c ij ) = c ij is the coefficient representing the connection relationship between sub-blocks based on the topological matrix C, x new is the online data, is the conditional probability of the abnormal operating state of the process.

[0164] The control limit corresponding to the i-th sub-block is the prior probability (1 - α) of the process occurring in an abnormal working condition:

[0165] Determine whether the monitoring statistic of the sub-block exceeds the control limit corresponding to the monitoring statistic. If so, it indicates that the current sub-block has a fault. If not, the operating state of the current sub-block is normal, and the CCA monitoring results of each sub-block are obtained.

[0166] Step S7: Calculate the overall statistical index FD of all sub-blocks, and determine whether there is a fault in the process by comparing the overall statistical index FD with the control limit, so as to realize the online monitoring of the operating state of the process.

[0167] The overall statistical index FD is used to represent the monitoring situation of the entire process. The expression of the overall statistical index FD is:

[0168]

[0169] where is the monitoring statistic of the B sub-blocks, and CL r,B is the control limit corresponding to the B-th sub-block, which is obtained by the Kernel Density Estimation (KDE) method. According to the definition of FD, its control limit can be obtained as 1. Determine whether the overall statistical index FD exceeds the obtained control limit. If so, it indicates that the operating state of the current process has a fault. If not, the operating state of the current process is normal.

[0170] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and gives the detailed implementation manner and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0171] In this embodiment, the distributed canonical correlation analysis process monitoring method based on partial sub-block communication proposed by the present invention is applied to the urban sewage treatment process, and the urban sewage treatment process is a general benchmark model BSM1 (Benchmark Simulation Model No. 1, BSM1).

[0172] The BSM1 benchmark simulation model is a sewage treatment process model proposed by the European Union's Science and Technology Cooperation Organization, and is widely used in the research of sewage treatment processes by personnel and institutions around the world.

[0173] In fact, the BSM1 model is mainly composed of two models, namely the activated sludge process model and the double exponential sedimentation model. The BSM1 benchmark simulation model mainly focuses on removing carbon and nitrogen in sewage, so that the treated wastewater can meet the secondary discharge standard.

[0174] Figure 3Disclosed is a flowchart of the Benchmark Simulation Model No. 1 according to an embodiment of the present invention, as Figure 3 shown. The process structure of BSM1 consists of a bioreactor and a secondary sedimentation tank.

[0175] The bioreactor consists of 5 well-mixed units. The first 2 units are anaerobic closed containers (referred to as anoxic tanks or anaerobic tanks), and the last 3 units are aerobic exposed containers (referred to as aeration tanks or aerobic tanks).

[0176] In this model, organic matter is mainly consumed through the vital activities of microorganisms, and nitrogen is removed through nitrification by microorganisms in the aerobic zone and denitrification in the anaerobic zone.

[0177] The total height of the secondary sedimentation tank is 4 meters, divided into 10 equal-height layers. And no chemical reactions occur in the secondary sedimentation tank, mainly physical filtration. The water in the reaction unit flows out from the fifth reaction tank. Part of it flows into the sixth layer of the secondary sedimentation tank, and the other part is recycled back to the inlet of the reaction unit through a pipeline to complete the internal cycle. In the secondary sedimentation tank, the wastewater meeting the sewage discharge standard after precipitation is discharged from the top of the secondary sedimentation tank into the river, and part of the sedimented sludge at the bottom is sent back to the first reaction tank for sludge recycling, and the other part is discharged for landfill.

[0178] Two faults are introduced into the BSM1 model, as shown in Table 1. The first fault is that the maximum growth rate of autotrophic bacteria in the first aerobic reaction tank undergoes a step change, decreasing from 0.5 to 0.3. Due to the weakened vital activities of autotrophic bacteria, it will affect the complex biochemical reactions in the urban sewage treatment process, thus changing the state of the entire system.

[0179] The second fault case is that the oxygen transfer coefficient of bioreactor 4 drops to half of the original value. The oxygen transfer coefficient has a great relationship with the dissolved oxygen concentration in the reaction tank, and the dissolved oxygen concentration is an important element in the urban sewage treatment process reaction, which is related to the sewage treatment effect. Because whether it is the microbial activities for removing organic matter or the denitrification of heterotrophic bacteria, they all require the participation of oxygen.

[0180] Table 1 Two faults of the BSM1 process

[0181] Fault Description Type 1 The maximum growth rate of autotrophic bacteria in the first aerobic reactor Step 2 The oxygen transfer coefficient of bioreactor 4 Step

[0182] Experimental simulation and application verification are carried out, and the specific process is as follows:

[0183] First, according to the sewage treatment process and structure of the BSM1 model, the 15 process measurement data for monitoring are divided into 3 sub-blocks. Sub-block 1 is the relevant water quality measurement variables of the treated effluent, Sub-block 2 is the water quality component measurement variables of the sewage influent and reaction tank 3, and Sub-block 3 is the measurement variables of reaction tanks 4 and 5. The measurement variables and sub-block divisions are shown in Tables 2 and 3 respectively:

[0184] Table 2 Measurement Variables of the BSM1 Process

[0185]

[0186]

[0187] Table 3 Sub-block Division of BSM1 Process Variables

[0188]

[0189] After dividing the process data into sub-blocks, the expression of its topological matrix C is determined according to process knowledge as:

[0190]

[0191] According to the urban sewage treatment process, the influent flows through two anaerobic tanks and three aerobic tanks respectively, removes organic matter and nitrogen, and then enters the secondary sedimentation tank. After clarification, the effluent is discharged into the river. Therefore, the topological coefficients between Sub-block 2 representing the influent and the third reaction tank and Sub-block 3 representing the fourth and fifth reaction tanks are set to 1, and the topological coefficient from Sub-block 3 to Sub-block 1 representing the effluent water quality components is also 1. Therefore, in this embodiment, 2 CCA models are established.

[0192] Set the number of canonical correlation components to 3, and the prior probability α of the normal operation state of the process to 99%.

[0193] As Figures 4 to 11 shown, the detection effects of two faults in the urban sewage treatment process of the present invention are demonstrated. Among them, Figure 4 and Figure 5 are the monitoring result diagrams of Sub-block 1 for faults 1 and 2 respectively, Figure 6 and Figure 7 are the monitoring result diagrams of Sub-block 2 for faults 1 and 2 respectively, Figure 8 and Figure 9 are the monitoring result diagrams of Sub-block 3 for faults 1 and 2 respectively, Figure 10 and Figure 11 are the monitoring result diagrams of the whole for faults 1 and 2 according to an embodiment of the present invention.

[0194] As shown in Table 4, the detection results of two faults in the BSM1 process by the proposed distributed CCA process monitoring method based on partly-connected sub-block communication (Partly-connected DCCA) are presented. From Table 4, it can be seen that for the second fault, the oxygen transfer coefficient of the 4th bioreactor decreased to half of the original value, and its fault detection rate was 100%, indicating that the present invention can well detect abnormalities;

[0195] For the first fault, a step change occurred in the maximum growth rate of autotrophic bacteria in the first aerobic reactor. The detection rates of the CCA method based on partly-connected sub-blocks and the proposed method of collaborative modeling between intra-blocks and inter-blocks were 59%. The reason for the analysis may be that the influence range of this fault is relatively wide.

[0196] Table 4 Fault detection rates of the BSM1 process by DCCA with partly-connected sub-block communication

[0197]

[0198] Overall, a distributed canonical correlation analysis process monitoring method and device based on partly-connected sub-block communication proposed by the present invention. While the CCA modeling based on partly-connected sub-block communication has good detection effects, it reduces the computational complexity of process monitoring and the load of data transmission, meets the requirements of process monitoring security, is more in line with the needs of actual production process monitoring, and has more engineering practical significance.

[0199] In addition, after detecting a fault, the distributed canonical correlation analysis process monitoring method based on partly-connected sub-block communication can also preliminarily locate the scope of the fault occurrence by checking the specific detection conditions of each sub-block.

[0200] For large-scale production processes with multiple connection units, the present invention adopts a multi-block statistical modeling method for distributed monitoring. First, it realizes process decomposition based on process knowledge and proposes a topological matrix to represent the connection situation between sub-blocks. Then, CCA modeling is carried out between sub-blocks with information transfer, and a residual generator is constructed according to the correlation relationship between sub-blocks. Next, a statistic is constructed through the residual vector to monitor the process.

[0201] The residual generator is used to avoid directly transmitting the original data, thereby reducing the amount of data communication, further improving the security of data transmission. At the same time, by means of block division, establishing a local model can better detect small faults, improve the sensitivity of monitoring, and contribute to the location of faults.

[0202] In this embodiment, the effectiveness of a distributed canonical correlation analysis process monitoring method and device based on partly-connected sub-block communication proposed by the present invention is verified through the BSM1 process simulation.

[0203] The present invention proposes a distributed canonical correlation analysis process monitoring device based on partial sub-block communication. The distributed canonical correlation analysis process monitoring device based on partial sub-block communication may include an internal communication bus, a processor, a read-only memory (ROM), a random access memory (RAM), a communication port, and a hard disk. The internal communication bus can enable data communication among the components of the distributed canonical correlation analysis process monitoring device based on partial sub-block communication. The processor can make judgments and issue prompts. In some embodiments, the processor may consist of one or more processors.

[0204] The communication port can enable data transmission and communication between the distributed canonical correlation analysis process monitoring device based on partial sub-block communication and external input / output devices. In some embodiments, the distributed canonical correlation analysis process monitoring device based on partial sub-block communication can send and receive information and data from a network through the communication port. In some embodiments, the distributed canonical correlation analysis process monitoring device based on partial sub-block communication can perform data transmission and communication with external input / output devices in a wired form through the input / output terminal.

[0205] The distributed canonical correlation analysis process monitoring device based on partial sub-block communication may further include different forms of program storage units and data storage units, such as a hard disk, a read-only memory (ROM), and a random access memory (RAM), which can store various data files used for computer processing and / or communication, as well as possible program instructions executed by the processor. The processor executes these instructions to implement the main part of the method. The results processed by the processor are transmitted to an external output device through the communication port and displayed on the user interface of the output device.

[0206] For example, the implementation process file of the above-mentioned distributed canonical correlation analysis process monitoring device based on partial sub-block communication can be a computer program, stored in the hard disk and can be recorded into the processor for execution to implement the method of the present invention.

[0207] When the implementation process file of the distributed canonical correlation analysis process monitoring method based on partial sub-block communication is a computer program, it can also be stored in a computer-readable storage medium as an article. For example, the computer-readable storage medium may include, but is not limited to, magnetic storage devices (e.g., hard disks, floppy disks, magnetic strips), optical disks (e.g., compact discs (CDs), digital versatile discs (DVDs)), smart cards, and flash memory devices (e.g., electrically erasable programmable read-only memories (EPROMs), cards, sticks, key drives). In addition, the various storage media described herein can represent one or more devices and / or other machine-readable media for storing information. The term "machine-readable medium" may include, but is not limited to, wireless channels and various other media (and / or storage media) that can store, contain, and / or carry code and / or instructions and / or data.

[0208] The distributed canonical correlation analysis process monitoring method and device based on partial sub-block communication provided by the present invention specifically have the following beneficial effects:

[0209] 1) A multi-block statistical modeling method is used for distributed process monitoring. The entire factory process needs to be decomposed into several different parts, which can improve the fault tolerance of the process. When there are problems with individual blocks, the remaining blocks can still be used for monitoring.

[0210] 2) The process is detected based on partial sub-block communication, which improves the sensitivity analysis of the process monitoring system. And if a fault is detected during the process, the fault location can be determined within a specific block or several blocks, and further diagnostic steps can be carried out to determine the root cause of the fault.

[0211] 3) The process data is divided into blocks based on process knowledge. By checking the specific detection conditions of each sub-block, the preliminary scope of the fault can be located after the fault is detected.

[0212] 4) The process is divided into individual sub-blocks for modeling respectively, which can effectively reduce the computational complexity of the monitoring process.

[0213] 5) Based on the topological matrix C used to describe the connection mode between sub-blocks, the partial sub-blocks of communication are modeled. While reducing the communication load of data transmission, a distributed CCA process monitoring that is more in line with the actual industrial process is established, meeting the requirements of process monitoring security, more in line with the needs of actual production process monitoring, and more practically significant in engineering.

[0214] Although the above methods are illustrated and described as a series of actions to simplify the explanation, it should be understood and appreciated that these methods are not limited by the order of the actions. Because according to one or more embodiments, some actions may occur in a different order and / or concurrently with other actions that are illustrated and described herein or that are not illustrated and described herein but are understood by those skilled in the art.

[0215] Those skilled in the art will appreciate that information, signals, and data can be represented using any of a variety of different technologies and techniques. For example, the data, instructions, commands, information, signals, bits, symbols, and chips described throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or magnetic particles, optical fields or optical particles, or any combination thereof.

[0216] Those skilled in the art will further appreciate that the various illustrative logical blocks, modules, circuits, and algorithmic steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability of hardware and software, the various illustrative components, blocks, modules, circuits, and steps are described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present invention.

[0217] The various illustrative logical modules and circuits described in connection with the embodiments disclosed herein can be implemented or performed with a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.

[0218] The steps of a method or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of both. The software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read from, and write to, the storage medium. In the alternative, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a user terminal.

[0219] As shown in this application and the claims, unless the context clearly indicates otherwise, words such as "a", "an", "one", and / or "the" are not specifically singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of the steps and elements that have been clearly identified, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.

[0220] The above embodiments are provided for those skilled in the art to implement or use the present invention. Those skilled in the art can make various modifications or changes to the above embodiments without departing from the inventive concept of the present invention. Therefore, the protection scope of the present invention is not limited by the above embodiments, but should be the maximum scope that conforms to the innovative features mentioned in the claims.

Claims

1. A distributed canonical correlation analysis process monitoring method, characterized in that It includes an offline modeling stage and an online monitoring stage. The offline modeling stage specifically includes the following steps: Step S1: Divide the training sample data into blocks and determine the topological matrix used to represent the connection situation between sub-blocks; Step S2: Based on the connection situation of the topological matrix, perform canonical correlation analysis modeling on the data of the current sub-block and the sub-blocks from the connections, extract the canonical correlation components and retain the projection matrix, and construct a residual generator according to the correlation relationship between sub-blocks; Step S3: Generate a residual vector according to the residual generator and calculate the statistic; The online monitoring stage specifically includes the following steps: Step S4: Perform corresponding block division on the new sampling data according to the method in the offline training stage; Step S5: Substitute the block-divided sampling data obtained in Step S4 into the corresponding canonical correlation analysis model respectively, and obtain the canonical correlation components of the new sample data through transformation by the projection matrix; Step S6: Perform fault detection on each sub-block and obtain the operating state of the sub-block; Step S7: Calculate the overall statistical index FD of all sub-blocks, and judge whether there is a fault in the process by comparing the overall statistical index FD with the control limit, so as to realize the online monitoring of the operating state of the process; Among them, Step S6 further includes the following steps: Step S601: The current sub-block receives the canonical correlation components from the connected sub-blocks and generates a residual vector; Step S602: Calculate the value of the statistic according to the residual vector. When there is one connected sub-block for the same sub-block, compare the statistic with the control limit; When there are multiple connected sub-blocks for the same sub-block, obtain multiple statistics, use the decision fusion method to fuse the multiple statistics of the same sub-block to obtain the monitoring statistic, and judge whether the value of the monitoring statistic of the sub-block exceeds the threshold. If so, it means that the current sub-block has a fault. If not, the operating state of the current sub-block is normal, and the canonical correlation analysis monitoring results of each sub-block are obtained.

2. The distributed canonical correlation analysis process monitoring method according to claim 1, characterized in that Step S1 further includes: Use a block division method based on process knowledge to divide the training sample data. According to the connection situation of the device equipment and operation units in the actual industrial production process, use prior knowledge and expert experience to divide the process.

3. The distributed canonical correlation analysis process monitoring method according to claim 1, characterized in that In Step S1, the topological matrix C is used to describe the connection relationship between sub-blocks. Among them, each element represents whether there is information interaction between the corresponding two sub-blocks. If so, it is set to 1. If not, it is 0.

4. The distributed canonical correlation analysis process monitoring method according to claim 1, wherein In Step S2, constructing the residual generator specifically includes the following steps: Step S201: Perform canonical correlation analysis modeling on each pair of sub-blocks according to the topological matrix C. Take sub-blocks X1 and X2 as the input and output variables of the canonical correlation analysis algorithm respectively, and extract the canonical correlation components of sub-block X1 and sub-block X2 respectively. The expressions of the canonical correlation components of sub-block X1 and sub-block X2 are: u 12 = A 12 T X1 v 12 = B 12 T X2 Among them, A 12 and B 12 are the projection matrices of two sub-blocks respectively; Step S202: Based on the canonical correlation components of each pair of sub-blocks, construct a residual generator. The expression of the residual generator between sub-blocks X1 and X2 is: r 12 = B 12 T X2 - Λ k,12 A 12 T X1 Among them, Λ k,12 is the canonical correlation coefficient diagonal matrix of the first k pairs of relevant variables of sub-block X1 and sub-block X2, and the value of k is obtained by cross-validation.

5. The distributed canonical correlation analysis process monitoring method according to claim 4, wherein In Step S201 described above, perform canonical correlation analysis modeling on each pair of sub-blocks according to the topological matrix C: When a sub-block has multiple connected sub-blocks, for the i-th sub-block, construct canonical correlation analysis models; where c ij is an element of the topological matrix C, and B is the number of sub-blocks.

6. The distributed canonical correlation analysis process monitoring method according to claim 1, characterized in that The statistic in Step S3 described above has the corresponding expression: Among them, is the statistic of sub-block X1, n is the degree of freedom, and r 12 is the residual vector between sub-block X1 and sub-block X2. Correspondingly, the statistic of sub-block X2 is calculated by swapping the input and output in the residual generator.

7. The distributed canonical correlation analysis process monitoring method according to claim 1, characterized in that In the step S602 described above, Bayesian inference is selected as the decision fusion method, and the expression of the monitoring statistic of the i-th sub-block is: Among them, is the monitoring statistic of the i-th sub-block, P(c ij ) = c ij is the coefficient representing the connection relationship between sub-blocks based on the topological matrix C, x new is the online data, is the conditional probability of the abnormal operating state of the process, c ij is the element of the topological matrix C.

8. The distributed canonical correlation analysis process monitoring method according to claim 1, wherein In the step S602 described above, the control limit corresponding to the i-th sub-block is the prior probability (1-α) of the process occurring in an abnormal working condition: Judge whether the monitoring statistic of the sub-block exceeds the control limit corresponding to the monitoring statistic. If so, it is considered that the sub-block has a fault. If not, it is considered that the sub-block is operating under normal working conditions.

9. The distributed canonical correlation analysis process monitoring method according to claim 1, wherein The overall statistical index FD in the step S7 described above has the corresponding expression: Among them, is the monitoring statistic of the B sub-blocks, CL r,B is the control limit corresponding to the Bth sub-block.

10. A distributed canonical correlation analysis process monitoring device, comprising: A memory for storing instructions executable by a processor; A processor for executing the instructions to implement the method according to any one of claims 1-9.

11. A computer-readable medium having computer instructions stored thereon, wherein when the computer instructions are executed by a processor, the method according to any one of claims 1-9 is executed.

Citation Information

Patent Citations

  • Distributed process monitoring method and device for intra-block and inter-block collaborative modeling

    CN116068974A