A wastewater treatment fault monitoring method based on manifold learning

By applying manifold learning and support vector data description methods in wastewater treatment, low-dimensional features of high-dimensional data are extracted and fault monitoring models are established, the problem of difficulty in monitoring nonlinear and non-Gaussian data in the prior art is solved, and a more efficient fault monitoring effect is achieved.

CN115310529BActive Publication Date: 2025-06-17SHANDONG HUATAI PAPER CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210924294.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-02
Publication Date
2025-06-17
Estimated Expiration
2042-08-02

AI Technical Summary

Technical Problem

The existing fault monitoring methods during wastewater treatment are difficult to effectively monitor nonlinear and non-Gaussian data, resulting in "dimensional disaster" problems, affecting the identification of fault points and the performance of fault classification.

Method used

A manifold learning-based method is adopted to extract the low-dimensional features of high-dimensional data through uniform manifold approximation and projection algorithm, and a fault monitoring model is established in combination with the support vector data description method to realize fault monitoring during wastewater treatment.

Benefits of technology

It effectively reduces the complexity of the fault monitoring model, avoids the "dimensional disaster" problem, improves the accuracy and performance of fault monitoring, and can better adapt to nonlinear and non-Gaussian data in industrial processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115310529B_ABST
    Figure CN115310529B_ABST
Patent Text Reader

Abstract

The present invention discloses a wastewater treatment fault monitoring method based on manifold learning, which can be used for fault monitoring of industrial processes with strong nonlinearity and non-Gaussianity. First, the manifold learning method is used to discover the intrinsic relationship between input data in the high-dimensional space, and the gradient descent method is used to find the low-dimensional feature data closest to the high-dimensional space data. With the help of the uniform manifold approximation and projection algorithm, the dimension of the process data is reduced, and the support vector data description classification algorithm is applied to quickly and effectively classify the low-dimensional feature data to solve the non-Gaussianity problem of data in the actual production process; the fault data of the wastewater treatment process is used to verify the fault monitoring effect of the model. The experimental results show that the uniform manifold approximation and projection combined with the support vector data description model can improve the monitoring effect of the fault monitoring model and is more suitable for the process monitoring system of complex industrial processes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of fault monitoring in the wastewater treatment process, and particularly relates to a wastewater treatment fault monitoring method based on manifold learning. Background Art

[0002] With the continuous development of modern industry, the production process has gradually tended to be continuous and large-scale. Therefore, there are relatively high requirements for the monitoring of quality indicators in the industrial process. In order to avoid losses caused by faults in the wastewater treatment process, it is necessary to monitor the faults in the wastewater treatment process in a timely and accurate manner. The process monitoring model can monitor the production system in real time, detect abnormalities in time before the production system is abnormal or the sensor fails, and judge the key variables causing the abnormalities, providing guidance for subsequent fault diagnosis and fault recovery. Currently, the fault monitoring method for the wastewater treatment process is as follows: using the historical process data in the normal operation state to establish a fault monitoring model, calculating the statistical control limit of the fault monitoring model, and then using the fault monitoring model to verify the real-time data of the new sample points in the wastewater treatment process monitored in real time. If the statistical control index of the new sample points exceeds the statistical control limit, the new sample points are determined as fault points.

[0003] In order to achieve fault monitoring in the wastewater treatment process, it is usually necessary to perform multivariate statistics on the relevant information of the sample points in the wastewater treatment process. The fault monitoring method for the wastewater treatment process based on multivariate statistics has several disadvantages. First, it is difficult to provide the direct effect of additional labels and can only be inferred through the projection results. Therefore, the selection of weights usually requires a repeated adjustment process. Second, due to the limitations of linear methods, it is necessary to perform low-dimensional mapping by maximizing the variance of the given data set that can be explained. Therefore, some data samples with a wider distribution may dominate the projection. This usually leads to the problem of "curse of dimensionality" (for example, normal samples and early-occurring fault samples overlap in the feature space). Especially when multiple faults need to be considered, the problem of "curse of dimensionality" will become more prominent. This problem not only hinders the identification of fault points by the fault monitoring model during the monitoring process, but also reduces the discrimination performance of the classifier for fault classification based on the extracted features. When the data feature dimension increases, the useful information does not increase while the noise continues to accumulate, the model complexity increases, thus affecting the performance of the classifier. Summary of the Invention

[0004] Manifold learning, full name Manifold Learning method, assumes that the data is uniformly sampled from a low-dimensional manifold in a high-dimensional Euclidean space. Manifold learning is to recover the low-dimensional manifold structure from the high-dimensional sampled data, that is, to find the low-dimensional manifold in the high-dimensional space and find the corresponding embedding mapping to achieve dimensionality reduction or data visualization.

[0005] To solve the above technical problems, the present invention proposes a wastewater treatment fault monitoring method based on manifold learning. Based on the uniform manifold approximation and projection (UMAP) algorithm and combined with the support vector data description (SVDD) method, the faults in the wastewater treatment process are monitored. First, the UMAP algorithm is used to extract the metric and non-metric attributes in the high-dimensional data and convert them into a common metric, that is, similarity, so as to utilize both labeled and unlabeled data simultaneously. The nonlinear structure of the dataset is extracted by projection while maintaining the topological structure, without the need for complex nonlinear kernel mapping. In addition to being able to identify the natural clusters in the dataset, the UMAP algorithm also provides the global structure information in the latent space, that is, the similarity and differences between clusters can be inferred according to their proximity in the latent space. The UMAP algorithm effectively considers the edge weights of the mapped graph in the low-dimensional space. Then, the SVDD method is used to establish a fault monitoring model for the low-dimensional feature data. After verification, the improved fault monitoring model has better fault monitoring effect compared with the traditional linear method.

[0006] Aiming at the nonlinear and non-Gaussian problems of industrial process data, the present invention provides a wastewater treatment fault monitoring method based on manifold learning. Based on the UMAP algorithm, both labeled and unlabeled data are utilized simultaneously. Based on the SVDD method, a fault monitoring model is established for the low-dimensional feature data to monitor the faults in the wastewater treatment process, which specifically includes the following steps:

[0007] S1. Collect the historical process data under normal operating conditions in the wastewater treatment process as the training set X train , and collect the real-time monitoring data containing fault features as the test set X test , and perform standardized preprocessing on the data of the training set X train and the test set X test to eliminate the influence caused by different dimensions among data variables. Standardize the data into standard data with a mean of 0 and a variance of 1.

[0008] Since the reliability of the model is verified using simulation data, the real-time monitoring data uses the sampled data under the fault state, and the fault information is known. By using the sample data containing fault information for verification, if it is not known which are the faults, the effect of the fault monitoring model cannot be judged.

[0009] S2. Determine the training set X through the maximum likelihood estimation methodtrain The intrinsic dimension is determined to find the number of characteristic variables in the low-dimensional space, and the cross-entropy in the high-dimensional space and the low-dimensional space is minimized using the gradient descent method.

[0010] S3. For the training set X train The data is dimensionally reduced to obtain the low-dimensional feature data Y train , and the least squares regression method is used to perform regression fitting on the original training set X train and the low-dimensional feature data Y train to obtain the mapping matrix A; the test set X test is projected into the low-dimensional space to obtain the test set sample y test .

[0011] S4. In the low-dimensional space, the support vector data description method is used to train the dimensionally reduced low-dimensional feature data Y train to obtain the relevant parameters of the hypersphere, overcoming the non-Gaussianity in the process data. The boundary function of the support vector data description method can be characterized by a hypersphere. By mapping the target sample points into a high-dimensional space feature space where spherical description is easier, the data is classified.

[0012] 5. For the data of the test set X test , the trained fault monitoring model is used to verify the monitoring effect of the fault monitoring method to determine whether the new sample points are fault points.

[0013] The data in step S1 is from the fault simulation data of the wastewater simulation benchmark platform. The training set includes the wastewater simulation benchmark model operating for 14 days under normal operating conditions, sampling once every 15 minutes on average, and accumulating 1345 samples. The test set contains the operating data of the wastewater simulation benchmark model under different fault modes. Perturbations are introduced into the simulation model on the 4th day to construct different types of faults to verify the fault detection model. That is, the fault is introduced starting from the 289th sample point, and perturbations are introduced into the simulation model to change the operating state of the model, thus obtaining the process data in the fault state.

[0014] The specific steps of step S2 are as follows:

[0015] S21: Assume that n high-dimensional space sample points X train ={x1, x2,..., x n}(x ∈ R m ), where m ≥ n.

[0016] For each data point x i define the local pseudo-metric space where X i is the set of k neighbors containing x i . The calculation formula is as follows:

[0017]

[0018] Among them, σ i and ρ i satisfy equations (2) and (3)

[0019]

[0020]

[0021] Among them, represents the Euclidean distance between points x i and x j . ρ i is the distance from point x i to its first nearest neighbor, ensuring that x i is locally connected to at least one of its nearest neighbors (i.e., the distance is 0), so that all sample points are not isolated. σ i is the diameter of the nearest neighbor data points of x i .

[0022] S22: Given a hyperparameter k, the set of neighbors of x i under d can be obtained as {x i1 ,..., x ik}, then the high-dimensional space fuzzy topological structure can be represented by an exponential probability distribution:

[0023]

[0024] S23: Symmetrize the exponential probability function:

[0025] P ij = P i|j + P j|i - P i|j P j|i (5)

[0026] S24: Assume that N mapped points Y train = {y1, y2,..., y N} in the low-dimensional space, y ∈ R n ), the probability distribution of the similarity between sample points in the low-dimensional space is as shown in equation (6):

[0027] q ij = (1 + a(y i - y j )) 2b ) -1 (6)

[0028] Among them, the hyperparameters a and b are determined by the fault monitoring model after optimal fitting according to the process data. All the data used in the modeling stage are from the training set, and the test set data is used to verify the monitoring effect of the fault monitoring model, that is, whether the faults can be detected.

[0029] S25: To make the similar sample points in the original high-dimensional space data remain as similar as possible in the projection space, the cross-entropy is introduced as the cost function, and the cross-entropy is expressed as:

[0030]

[0031] By minimizing the cost function, the low-dimensional space data with the highest fitting degree to the high-dimensional space data is obtained. The gradient descent method is used to optimize the cost function to obtain the minimum value of the cost function, and the result after uniform manifold approximation and projection dimensionality reduction is obtained. To improve the dimensionality reduction effect, the original high-dimensional space data can be iterated multiple times to improve the correctness of the mapping of the low-dimensional space data. Using the gradient descent method, after multiple iterations, the parameters when the cost function is minimized are obtained, and the low-dimensional feature matrix is obtained.

[0032] The specific steps of step S3 are as follows:

[0033] S31: By training a linear least squares regression model, the sample points of the test set are embedded into the low-dimensional space to obtain the projection matrix A. In the wastewater treatment process, considering n variables under normal conditions, each variable has m measured variables, that is, each observed value x i is an m-dimensional vector. Let X train =[x1, x2,..., x n ∈R m×n be the original training data with n samples, where x i ∈R m (i = 1, 2,..., n) represents the m variables measured by the sensor at the i-th sampling time. Y = [y1, y2,..., y n ∈R l×n is the data embedded in the l-dimensional space, where y i ∈R l (i = 1, 2,..., n) is the mapping point of x i in the l-dimensional space.

[0034] x i →y i = A T x i (8)

[0035] S32: Based on the assumption of the existence of potential local manifolds in the original high-dimensional space process data, the projection matrix is calculated using linear least squares regression:

[0036] A = (X train X train T ) -1 X train Y train T (9)

[0037] wherein, is the linear projection matrix of the mapping relationship between the high-dimensional space data and the low-dimensional space data.

[0038] The specific steps of S4 are as follows:

[0039] S41: Given a training sample X test = {x1, x2,..., x n}(x i ∈R m ), which constitutes n learning samples of the single-value classifier. The goal of the support vector data description is to find a hypersphere with the smallest volume that can contain all samples or the vast majority of samples as much as possible. Define a non-linear mapping ψ(·): x i →ψ(x i ), which maps the entire data set to the high-dimensional feature space. Then the optimization problem is expressed as:

[0040]

[0041] s.t. ||ψ(x i ) - a|| 2 ≤R 2 + ξ i (11)

[0042] wherein, a is the center of the hypersphere, R is the radius of the hypersphere, and ||ψ(x i ) - a|| is the distance from the sample point xi to the center a of the sphere. ξ i is the relaxation factor, and ξ i ≥0, and C is the penalty coefficient. Introducing these two parameters is to balance the hypersphere volume and misclassified samples to make the classification result more accurate.

[0043] S42: Introduce the Lagrange multiplier α = [α1, α2,..., α N T , and the above target optimization problem can be transformed into a Lagrangian extreme value problem:

[0044]

[0045] s.t. 0 ≤ α i ≤C (13)

[0046] ​

[0047] Among them, K(x i , x j ) = <ψ(x i ), ψ(x j )> is a kernel function. Using the Gaussian kernel function maps the original data space to a high-dimensional feature space, converts the non-linear problem in the original data into a linear problem in the high-dimensional space, and can be solved without knowing the specific form of the non-linear function.

[0048] The optimal solution is obtained through S42 satisfies In the formula, SV is the support vector. Calculate the center and radius of the hypersphere through formula (15) and formula (16).

[0049]

[0050]

[0051] Among them, x s represents the support vector.

[0052] The specific steps of the said step S5 are as follows:

[0053] S51: For the test set sample y test , calculate the distance from y test to the center of the hypersphere as:

[0054]

[0055] S52: Construct a fault monitoring model, and define the distance between the sample point and the center of the hypersphere as the corresponding monitoring statistic dis.

[0056] S53: Define the radius R of the hypersphere as the control limit DIS of the statistical monitoring index, and the logic of fault monitoring is:

[0057]

[0058] Advantages of the present invention:

[0059] The present invention combines a dimensionality reduction preprocessing method and a classifier model, designs a discrimination target on the classifier to label data, and thus realizes fault monitoring. The trained fault monitoring model is used to verify the monitoring effect of the fault monitoring method and determine whether a new sample point is a fault point. The advantage of this method is that, based on the support vector data description model, it combines the uniform manifold approximation and projection algorithm and the data preprocessing method to reduce the complexity of the fault monitoring model and avoid the problem of "dimensionality disaster" caused by too high dimensionality of process data. And because the uniform manifold approximation and projection algorithm can maximize the sample distance between different categories and minimize the sample distance between the same category during dimensionality reduction, it can extract the non-linear information in the process data. The support vector data description method can construct a hypersphere to contain the sample data of the same category to the greatest extent, and uses the support vector data description method to train the process data in the low-dimensional subspace to establish a fault monitoring model. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 is a flowchart of a fault monitoring model based on the uniform manifold approximation and projection algorithm.

[0061] Figure 2 is a projection diagram of the low-dimensional data preprocessed by the uniform manifold approximation and projection algorithm in a three-dimensional space.

[0062] Figure 3 is a fault monitoring diagram of the improved fault monitoring model. DETAILED DESCRIPTION OF THE INVENTION

[0063] The present invention will be further described more clearly and completely below. Obviously, the described examples are only a part of the examples of the present invention, rather than all the embodiments.

[0064] As Figures 1 - 3 shown, a wastewater treatment fault monitoring method based on manifold learning includes the following steps:

[0065] S1. Data preprocessing: Collect the historical process data under normal operating conditions during the wastewater treatment process as the training set X train , collect the real-time monitoring data containing fault characteristics as the test set X test , and standardize the data of the training set X train and the test set X test to eliminate the influence caused by different dimensions between process data variables.

[0066] The standardization formula is as follows:

[0067]

[0068] In the formula, X *Let the original data be \(X_0\), the standardized data be \(X\), \(\mu\) and \(\sigma\) be the mean and variance of all sample data respectively.

[0069] S2. Construct a uniform manifold approximation and projection model to extract the characteristic information of process data. By the maximum likelihood estimation method, determine the intrinsic dimension of the process data, determine the number of low-dimensional characteristic variables, and use the gradient descent method to minimize the cross-entropy in the high-dimensional space and the low-dimensional space.

[0070] S3. Perform dimensionality reduction on the training set \(X\) train to obtain low-dimensional characteristic data \(Y\) train , and use the least squares regression method to perform regression fitting on the original training set \(X\) train and the low-dimensional characteristic data \(Y\) train to obtain the mapping matrix \(A\); project the test set \(X\) test into the low-dimensional space to obtain the test set sample \(y\) test .

[0071] S4. Construct a support vector data description model. In the low-dimensional subspace, use the support vector data description method to train the dimensionality-reduced process data to obtain the relevant parameters of the hypersphere, overcome the non-Gaussianity in the process data, and establish a fault monitoring model.

[0072] S5. For the test set data, use the trained fault monitoring model to verify the monitoring effect of the fault monitoring method and determine whether the new sample point is a fault point.

[0073] Taking the experimental data of the BSM1 simulation benchmark platform as an example, based on the BSM1 simulation software, by changing the process variable parameters, 8 kinds of fault data can be generated: (1) the fault of the maximum specific growth rate change of autotrophic bacteria; (2) the fault of the maximum specific growth rate change of heterotrophic bacteria; (3) the fault of the sedimentation rate change of the secondary sedimentation tank; (4) the fault of the output of the nitrate actuator; (5) the change of the set value of the dissolved oxygen controller; (6 - 8) the drift fault, offset fault and complete failure fault of the dissolved oxygen sensor. 8 different fault modes are designed by changing the corresponding signal inputs. Combining Figure 1 The present invention is further described in detail as follows:

[0074] The process of establishing the model in the offline stage:

[0075] The first step: Collect the historical process data under normal operating conditions in the wastewater treatment process as the training set, and standardize the process data to eliminate the influence caused by different dimensions among the process data variables.

[0076] The second step: By the maximum likelihood estimation method, determine the intrinsic dimension of the process data, determine the number of low-dimensional characteristic variables, and use the gradient descent method to minimize the cross-entropy in the high-dimensional space and the low-dimensional space.

[0077] Step 3: Calculate the mapping coefficient matrix A after dimensionality reduction by the Isometric Mapping (ISOMAP) algorithm using the least squares regression method.

[0078] Step 4: In the low-dimensional subspace, use Support Vector Data Description (SVDD) to train the dimensionality-reduced process data y train to obtain the relevant parameters of the hypersphere, overcoming the non-Gaussianity in the process data.

[0079] Operation process of the online phase fault monitoring model:

[0080] Step 5: For Test data x new , calculate its mapping matrix y i = A T x new in the low-dimensional space to obtain the dimensionality-reduced test set data.

[0081] Step 6: Calculate the distance between the new data and the center of the hypersphere of the trained Support Vector Data Description model, i.e., the monitoring statistic dis.

[0082] Step 7: Determine whether the monitoring statistic dis of the new sample is greater than the control limit DIS of the monitoring index. If dis > DIS, the new sample point is regarded as a fault sample; otherwise, the new sample point is the process data under normal operating conditions. Table 1 shows the fault monitoring rates of the three fault monitoring models for 8 faults of the wastewater simulation benchmark model, and Table 2 shows the false alarm rates of the three fault monitoring models for 8 faults of the wastewater simulation benchmark model. By comparing the monitoring results, it can be seen that for faults 1, 2, 3, and 6, the improved fault monitoring model has better monitoring results, and for other fault types with poor monitoring effects, the monitoring effect has also been significantly improved. Faults 5 and 8 are offset types with relatively large changes, so they are relatively easy to be detected by the monitoring method. In the fault monitoring based on the improved model, the false alarm rate of fault 1 is significantly lower, but there is a certain amount of missed detection. The monitoring rates of faults 2 and 3 have increased significantly. For this kind of slowly changing fault, compared with the comparison algorithm, the monitoring delay is the smallest, and it can detect abnormalities in time at the early stage of the fault occurrence, give an early warning, and prevent the situation from deteriorating due to the further expansion of the fault. Fault 4 is an increase in the nitrate input amount, which belongs to an internal interference fault of the model. In the actual wastewater treatment process, due to the feedback control in the control loop, the offset fault signal is gradually corrected, and the fault information is not obvious. When using the Principal Component Analysis (PCA) method to reduce the dimensionality of the process data, traditional linear methods cannot capture the non-linear characteristics in the process data, resulting in low fault monitoring accuracy. However, when using the Isometric Mapping (ISOMAP) algorithm to reduce the dimensionality of the process data, it can better solve the non-linear problem of the wastewater treatment process data.

[0083] Table 1 Fault monitoring rate (%)

[0084]

[0085] Table 2 Fault misdetection rate (%)

[0086]

[0087] Considering the non - linearity and non - Gaussianity of the data in the wastewater treatment process and the uncertainty existing in the industrial process, it is difficult for the model in the fault monitoring process to achieve a better monitoring effect. The method of the present invention uses the uniform manifold approximation and projection algorithm to better explain the non - linearity of the data, and combines the support vector data description method to classify the fault data and normal data, so that the improved fault monitoring model can better adapt to the actual industrial process.

[0088] The process of evaluating the model effect of the fault monitoring model is specifically as follows: Bring the test set data into the trained fault monitoring model, and use the fault monitoring rate and the fault misdetection rate to comprehensively evaluate the monitoring effect of the fault monitoring model.

[0089] The above describes the basic principle, main features and the advantages of the present invention. The above - mentioned is only the preferred specific embodiment of the present invention, and the protection scope of the present invention is not limited thereto. Those skilled in the art can easily think of changes or substitutions within the technical scope shown by the present invention, and all of them should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be defined by the appended claims and their equivalents.

Claims

1. A wastewater treatment fault monitoring method based on manifold learning, characterized in that, Using both labeled and unlabeled data based on the Isometric Feature Mapping (ISOMAP) algorithm, a fault monitoring model is established for low-dimensional feature data based on the Support Vector Data Description (SVDD) method to monitor faults in the wastewater treatment process. The specific steps are as follows: S1. Collect historical process data under normal operating conditions during the wastewater treatment process as the training set X train , and collect real-time monitoring data containing fault characteristics as the test set X test , and perform standardized preprocessing on the data of the training set X train and the test set X test . Standardize the data into standard data with a mean of 0 and a variance of 1, and eliminate the influence caused by different dimensions among data variables; S2. Determine the training set X by the maximum likelihood estimation method train Determine the intrinsic dimension of, determine the number of feature variables in the low-dimensional space, and use the gradient descent method to minimize the cross-entropy in the high-dimensional space and the low-dimensional space; the specific steps are as follows: Define a local pseudo-metric space for each data point of the high-dimensional space sample points, represent the high-dimensional space fuzzy topological structure using an exponential probability function, symmetrize the exponential probability function, represent the similarity between sample points in the low-dimensional space using a similarity probability function, and introduce cross-entropy as a cost function; S3. Perform dimensionality reduction on the training set X train to obtain low-dimensional feature data Y train , and use the least squares regression method to perform regression fitting on the original training set X train and the low-dimensional feature data Y train to obtain the mapping matrix A; project the test set X test onto the low-dimensional space to obtain the test set sample y test ; S4. In the low-dimensional space, use the support vector data description method to train the low-dimensional feature data Y after dimensionality reduction train to obtain the relevant parameters of the hypersphere and overcome the non-Gaussianity in the process data; S5. For the data in the test set X test Use the trained fault monitoring model to verify the monitoring effect of the fault monitoring method and determine whether the new sample point is a fault point.

2. The wastewater treatment fault monitoring method based on manifold learning according to claim 1, characterized in that, The data in step S1 is sourced from the fault simulation data of the wastewater simulation benchmark platform. The training set includes the operation data of the wastewater simulation benchmark model for 14 days under normal conditions, with samples taken every 15 minutes on average, accumulating 1345 samples. The test set contains the operation data of the wastewater simulation benchmark model under different fault modes. A perturbation is introduced into the simulation model on the 4th day, that is, the fault is introduced starting from the 289th sample point.

3. The wastewater treatment fault monitoring method based on manifold learning according to claim 1, characterized in that, The specific steps of step S2 are as follows: S21: Assume n high-dimensional space sample points X train ={x1, x2, …, x n}(x ∈ R m ), where m ≥ n; For each data point x i Define a local pseudo - metric space where X i is the set containing the k neighbors of x i , and the calculation formula is: The calculation formula is: where σ i and ρ i satisfy equations (2) and (3) Among them, represents the Euclidean distance between point x i and x j , ρ i is the distance from point x i to its first nearest neighbor, ensuring that x i is locally connected to at least one of its nearest neighbors so that all sample points are not isolated, and σ i is the diameter of the nearest neighbor data points of x i ; S22: Given a hyperparameter k, obtain x i The set of neighbors of x under d {x i1 , …, x ik}, then the high-dimensional space fuzzy topological structure is represented by an exponential probability function: S23: Symmetrize the exponential probability function: P ij = P i|j + P j|i - P i|j P j|i (5) S24: Assume the mapped points Y in N low-dimensional spaces train ={y1, y2, …, y N}(y ∈ R n ). The similarity probability function between sample points in the low-dimensional space is as shown in Equation (6): q ij = (1 + a(y i - y j )) 2b ) -1 (6) Among them, the hyperparameters a and b are determined by the fault monitoring model through optimal fitting based on the process data; S25: In order to make similar sample points in the original high-dimensional space data remain as similar as possible in the projection space, cross-entropy is introduced as the cost function, and the cross-entropy is expressed as:

4. The wastewater treatment fault monitoring method based on manifold learning according to claim 1, characterized in that, The specific steps of step S3 are as follows: S31: By training a linear least squares regression model, the sample points of the test set are embedded into a low-dimensional space to obtain a projection matrix A; during the wastewater treatment process, considering n variables under normal conditions, each variable has m measurement variables, that is, each observed value x i is an m-dimensional vector; let X train =[x1, x2, …, x n ∈ R m×n be the original training data with n samples, where x i ∈ R m (i = 1, 2, …, n) represents the m variables measured by the sensor at the i-th sampling time; Y = [y1, y2, …, y n ∈ R l×n be the data embedded in the l-dimensional space, where y i ∈ R l (i = 1, 2, …, n) is the mapped point of x i in the l-dimensional space; x i →y i =A T x i (8) S32: Based on the assumption of the existence of potential local manifolds in the original high-dimensional space process data, use linear least squares regression to calculate the projection matrix: A = (X train X train T ) -1 X train Y train T (9) Among them, is a linear projection matrix of the mapping relationship between high-dimensional space data and low-dimensional space data.

5. The wastewater treatment fault monitoring method based on manifold learning according to claim 1, characterized in that, The specific steps of step S4 are as follows: S41: Given a training sample X test ={x1, x2, …, x n}(x i ∈R m ), which constitutes n learning samples of the single - value classifier. The goal of the support vector data description is to find a hypersphere with the smallest volume that can contain all samples or the vast majority of samples as much as possible; Define a non - linear mapping ψ(·): x i →ψ(x i ), mapping the entire data set to a high - dimensional feature space. Then the optimization problem is expressed as: Among them, a is the center of the hypersphere, R is the radius of the hypersphere, and ||ψ(x i ) - a|| is the distance from the sample point xi to the center a of the sphere; ξ i is the relaxation factor, and ξ i ≥ 0, and C is the penalty coefficient; S42: Introduce the Lagrange multiplier α = [α1, α2, …, α N T , and the above objective optimization problem can be transformed into a problem of finding the Lagrangian extreme value:​ Among them, K(x i , x j ) = <ψ(x i ), ψ(x j )> is a kernel function. Using the Gaussian kernel function maps the original data space to a high-dimensional feature space, converts the non-linear problem in the original data into a linear problem in the high-dimensional space, and can be solved without knowing the specific form of the non-linear function; Obtain the optimal solution through S42 Meet where SV is the support vector; calculate the center and radius of the hypersphere through equations (15) and (16); Among them, x s represents a support vector.

6. A wastewater treatment fault monitoring method based on manifold learning according to claim 1, wherein, The specific steps of step S5 are as follows: S51: For the test set sample y test , calculate the distance from y test to the center of the hypersphere as follows: S52: Construct a fault monitoring model, and define the distance between the sample point and the center of the hypersphere as the corresponding monitoring statistic dis; S53: Define the radius of the hypersphere as the control limit DIS of the statistical monitoring index. The logic of fault monitoring is:

7. A wastewater treatment fault monitoring method based on manifold learning according to any one of claims 1 to 6, wherein, The process of evaluating the model effect of the fault monitoring model is specifically as follows: Bring the test set data into the trained fault monitoring model, and use the fault detection rate and the false alarm rate to comprehensively evaluate the monitoring effect of the fault monitoring model.

Citation Information

Patent Citations

  • Batch process fault monitoring method based on multi-stage ICA-SVDD

    CN107272655A

  • Multi-modal chemical process fault detection method based on improved t-SNE

    CN113741364A