Method for monitoring fluctuations in a wastewater treatment process based on principal component analysis
By establishing a fluctuation model of the wastewater treatment process through principal component analysis, and using T2 and SPE indicators for real-time monitoring and early warning, the problem of difficulty in monitoring fluctuations during wastewater treatment is solved, and the stable operation of the system and the guarantee of effluent quality are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-31
- Publication Date
- 2026-06-30
Smart Images

Figure CN122310174A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of wastewater treatment monitoring technology, and specifically to a method for monitoring fluctuations in wastewater treatment processes based on principal component analysis (PCA). Background Technology
[0002] The activated sludge process is currently the mainstream wastewater treatment process. Fluctuations and shocks in wastewater treatment are widespread, potentially originating from upstream equipment fluctuations, environmental fluctuations, and fluctuations in microbial growth processes. These fluctuations are characterized by varying time scales, difficulty in prediction and control, reliance on passive monitoring and response, and unpredictable consequences. Problems arising from these fluctuations include: reduced pollutant treatment capacity, fluctuations in effluent pollutant indicators, increased operating costs, and increased difficulty in monitoring and response.
[0003] Monitoring fluctuations and shocks in the wastewater treatment process is a systematic task aimed at timely detection of problems, precise source identification, and effective response, thereby ensuring the stable operation of the treatment system and compliance with effluent quality standards. According to research, there is currently no effective systematic method for monitoring fluctuations in the wastewater treatment process. Summary of the Invention
[0004] The present invention aims to at least solve one of the technical problems mentioned in the background art above, and provides a method for monitoring fluctuations in wastewater treatment processes based on principal component analysis, so as to effectively monitor abnormal fluctuations in wastewater treatment processes.
[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0006] A method for monitoring fluctuations in wastewater treatment processes based on principal component analysis includes the following steps:
[0007] S1, collect various variables of the reaction process operation during historical wastewater treatment processes as historical variable data;
[0008] S2, Based on the collected historical variable data, run the PCA algorithm for offline modeling to establish a PCA normal fluctuation model, and calculate the statistical index T during the establishment of the PCA normal fluctuation model. 2 The threshold of SPE, the statistical indicator T 2 The SPE (Statistical Parameters) represents the distance between the projection of the variable data in the principal component space and the model center of the PCA normal fluctuation model.
[0009] S3 collects various variables related to the operation of the reaction process during wastewater treatment as real-time variable data;
[0010] S4, input the real-time variable data into the established PCA normal fluctuation model, and calculate the statistical index T. 2 SPE, and the calculated statistical index T 2 SPE and statistical indicator T 2 The threshold of SPE is compared, and if the calculated statistical index T is... 2 If the SPE value exceeds the corresponding threshold, an early warning will be issued.
[0011] Furthermore, in steps S1 and S3, data cleaning is performed on the collected historical and real-time variable data. The data cleaning method is as follows: the fluctuations of the variable data in the sewage treatment process are quantified at different time scales to obtain fluctuation quantification values; after completing the fluctuation quantification of the variable data, the time period of normal fluctuation is selected based on the fluctuation quantification values of the variable data; after completing the selection of the time period of normal fluctuation, the fluctuation quantification values of the variable data under different sampling frequencies are timestamped and matched to construct a matrix X, and the fluctuation quantification values are normalized accordingly to form a standard probability distribution with a mean of 0 and a variance of 1.
[0012] Furthermore, the method for quantifying the fluctuations of variable data in the wastewater treatment process across different time scales is as follows: divide the standard deviation of the variable data over a period of time by the mean over that period to obtain the quantified fluctuation value.
[0013] Furthermore, the steps for establishing the PCA normal fluctuation model include:
[0014] S21, Calculate the covariance matrix of the matrix X;
[0015] S22, perform eigenvalue decomposition on the covariance matrix and arrange the eigenvalues in descending order;
[0016] S23, calculate the percentage of information captured by each principal component based on the eigenvalues, and select the number of principal components based on this to determine the loading matrix, which is the core of the PCA model;
[0017] S24, Calculate statistical index T based on the loading matrix. 2 The threshold corresponding to SPE.
[0018] Further, in step S21, the covariance matrix of the matrix is calculated using the following formula;
[0019] ;
[0020] In the formula, matrix C is the covariance matrix, and N is the number of time points in matrix X.
[0021] Further, in step S22, the covariance matrix is decomposed into eigenvalues according to the following formula:
[0022] ;
[0023] In the formula, , The eigenvalues are eigenvalues; m is the number of columns in matrix X, which is the number of variable fluctuation values; the characteristic matrix P consists of the eigenvalues. The corresponding feature vector composition:
[0024] and , Let be the i-th element in the characteristic matrix.
[0025] Furthermore, in step S23, due to the eigenvalues This describes the information that each principal component can capture from the variable dataset; therefore, the percentage of information that each principal component can capture is calculated using the following formula:
[0026] ;
[0027] In the formula, PC i Let i be the i-th principal element;
[0028] because The data has already been sorted in descending order, so the percentage of total information captured by the first k principal elements is:
[0029] ;
[0030] The method for selecting the number of principal components is as follows: If the total information captured by the first k principal components is... Th is a predetermined information threshold, representing the boundary between information and noise; then k is taken as the number of principal components;
[0031] After determining the number of principal components, the loading matrix V of the principal component space can be determined based on the feature matrix P. The loading matrix V is the first k columns of data from the feature matrix P. .
[0032] Furthermore, in step S24, the statistical indicator T 2 threshold The calculation method is as follows:
[0033] ;
[0034] In the formula, m is the number of columns in matrix X, representing the number of variables in matrix X. It follows an F-distribution with the first degree of freedom k and the second degree of freedom mk, therefore Based on confidence level A function for calculating the confidence space;
[0035] Threshold of the statistical indicator SPE The calculation method is as follows:
[0036] ;
[0037] ;
[0038] In the formula, Let be the quantiles on the standard normal distribution corresponding to the desired confidence level α, and m be the number of columns in matrix X.
[0039] Furthermore, in step S4, the statistical indicator T 2 The calculation method is as follows:
[0040] Represent the current real-time variable data x as a row vector, and calculate the score of the current real-time variable data x for each principal component. Then, based on the score, calculate the corresponding real-time variable data x. As shown in the following formula:
[0041] ;
[0042] ;
[0043] The statistical indicator SPE is calculated as follows:
[0044] Calculate the error of reconstructing the current real-time variable data x. That is, the information difference between the real-time variable data x after reconstructing it through the principal component space and the real-time variable data x, and the SPE is calculated based on this, as shown in the following formula:
[0045] ;
[0046] .
[0047] Furthermore, in step S4, when the calculated statistical index T... 2 When SPE is greater than the corresponding threshold, the algorithm also calculates the relationship between each variable in the real-time variable data x and T. 2 The contribution of SPE is used to determine the source of volatility, thus completing the volatility origin tracing:
[0048] Variables versus statistical indicators T 2 The contribution is calculated as follows:
[0049] The j-th variable in the real-time variable data vector x For the i-th score Contribution for:
[0050] ;
[0051] In the formula To load the elements of matrix V at (i, j);
[0052] Then, by summing, the j-th variable is calculated. For T 2 Contribution for:
[0053] ;
[0054] Contribution of variables to the statistical indicator SPE The calculation method is as follows:
[0055] ;
[0056] In the formula Let be the absolute value of the error e.
[0057] By adopting the above technical solution, the present invention has the following beneficial effects:
[0058] The aforementioned method for monitoring fluctuations in wastewater treatment processes based on principal component analysis (PCA) addresses common fluctuations and shocks affecting the stable operation of wastewater treatment processes. This method utilizes various variables reflecting the operational status of the wastewater treatment process to establish an offline PCA normal fluctuation model, which is then deployed online to monitor fluctuations in these variables. When fluctuations exceed acceptable limits, an early warning is issued, effectively monitoring abnormal fluctuations in the wastewater treatment process.
[0059] The aforementioned wastewater treatment process fluctuation monitoring method based on principal component analysis, when fluctuations exceed the standard, also calculates the relationship between each variable and T in the real-time variable data. 2 The contribution of SPE is used to determine the source of fluctuations, complete the fluctuation source tracing, find the causal variables that cause the fluctuation data to exceed the limit, and guide operators to deal with fluctuations normally, thereby reducing the further impact of fluctuations and shocks and achieving the goal of stabilizing the wastewater treatment process. Attached Figure Description
[0060] Figure 1 This is a flowchart of a preferred embodiment of the wastewater treatment process fluctuation monitoring method based on principal component analysis according to the present invention.
[0061] Figure 2This diagram illustrates the timestamp format for variable data in the wastewater treatment process fluctuation monitoring method based on principal component analysis, which is a preferred embodiment of the present invention.
[0062] Figure 3 This diagram illustrates the PCA principal component number determination method for a preferred embodiment of the wastewater treatment process fluctuation monitoring method based on principal component analysis of the present invention. Figure 3 The table on the left shows the table based on each characteristic value. The calculation includes the percentage of information captured by each principal component and the percentage of total information captured by the first k principal components. Figure 3 The right figure shows the number of principal components and eigenvalues. The relationship diagram is shown, with the number of principal elements on the coordinate axis and the eigenvalues on the vertical axis.
[0063] Figure 4 In the PCA normal fluctuation model of the wastewater treatment process fluctuation monitoring method based on principal component analysis, which is a preferred embodiment of the present invention, T 2 The SPE statistical index and its threshold are defined, where the space formed by the purple cylinder is the principal component space. A coordinate system is established with the center of the bottom circle of the cylinder as the origin, and the principal component directions PC1 and PC2 (PCn in multidimensional space) as the coordinates. The height of the top circle of the cylinder is the SPE threshold, and the edge of the bottom circle is T. 2 Threshold; purple dots represent variable parameters within the threshold range, green dots represent SPE values outside the SPE threshold range, and yellow dots represent values within the T threshold range. 2 T outside the threshold range 2 value.
[0064] Figure 5 At a certain moment T, the wastewater treatment process fluctuation monitoring method based on principal component analysis, which is a preferred embodiment of the present invention, is shown. 2 The contribution plot has the horizontal axis representing the number of variables and the vertical axis representing the T value corresponding to each variable. 2 The contribution value of the graph. Detailed Implementation
[0065] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0066] In the description of this invention, it should be understood that the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0067] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0068] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0069] The accompanying drawings illustrate various structural schematic diagrams according to embodiments disclosed in this invention. These drawings are not to scale, and some details have been enlarged for clarity, and some details may have been omitted. The shapes of the various regions and layers shown in the drawings, as well as their relative sizes and positional relationships, are merely exemplary and may deviate from reality due to manufacturing tolerances or technical limitations. Furthermore, those skilled in the art can design regions / layers with different shapes, sizes, and relative positions as needed.
[0070] Please see Figure 1 A preferred embodiment of the present invention provides a method for monitoring fluctuations in a wastewater treatment process based on principal component analysis, comprising the following steps:
[0071] S1 collects various variables related to the operation of the reaction process during historical wastewater treatment, which are used as historical variable data.
[0072] In this embodiment, the variable data includes, but is not limited to, the following variables: aeration rate, online COD detection data, online total nitrogen detection data, online ammonia nitrogen detection data, pH detection data, wastewater return flow rate, inlet wastewater flow rate detection data, and wastewater treatment device outlet discharge detection data. In this embodiment, after obtaining the historical variable data, data cleaning is performed on the historical variable data to match the modeling requirements of PCA. The data cleaning method is as follows:
[0073] First, the fluctuations of variable data in the wastewater treatment process are quantified across different time scales to obtain quantified fluctuation values. Because the PCA method is used to monitor the fluctuation characteristics of key variables, these fluctuations need to be quantified. Specifically, the standard deviation of the variable data over a period of time is divided by the mean over that period to obtain the quantified fluctuation value. For example, the quantified fluctuation value for COD measured within one day is:
[0074] ;
[0075] in, This represents the daily fluctuation quantification value of the COD variable. This represents the daily average of the COD variable. Let d represent the daily standard deviation of the COD variable, where d represents the data at the daily time scale. The fluctuations of each variable in the wastewater treatment process can be quantified at different time scales. For example, COD fluctuations can be quantified at time scales such as hours (h), days (d), weeks (w), and months (m). Then, the quantified fluctuation values of different variables at different time scales can be collected uniformly as a dataset for PCA modeling, used to monitor the fluctuations of different variables at different time scales.
[0076] Secondly, after quantifying the volatility of the variable data, the time period of normal volatility is selected based on the quantified volatility values. PCA typically uses data with normal volatility to build the model, constructing a linear relationship model between the quantified volatility values of different variables at different time scales.
[0077] Finally, after selecting the time period for normal fluctuations, the quantized fluctuation values of variable data at different sampling frequencies are timestamped and matched, according to... Figure 2 The format diagram shown uses the quantized fluctuation values of multiple variable data arranged horizontally as rows of a matrix, and the quantized fluctuation values of the same variable data obtained at different sampling frequencies arranged vertically as columns of a matrix to construct matrix X. The quantized fluctuation values are then normalized to form a standard probability distribution with a mean of 0 and a variance of 1.
[0078] After selecting the normal data time period, it is necessary to align and match the timestamps of data from different sampling frequencies in order to perform modeling using the PCA method. In this implementation, it is required that different variables be aligned according to... Figure 2The format is organized by arranging the quantized fluctuation values of multiple variables horizontally as rows of matrix X, and arranging the quantized fluctuation values of the same variable obtained at different sampling times (i.e., timestamps) vertically as columns of matrix X to construct matrix X. The quantized fluctuation values are then normalized to form a standard probability distribution with a mean of 0 and a variance of 1, in order to eliminate the interference of different variable fluctuation amplitudes and dimensions on the PCA algorithm.
[0079] S2, Based on the collected historical variable data, run the PCA algorithm for offline modeling to establish a PCA normal fluctuation model, and calculate the statistical index T during the establishment of the PCA normal fluctuation model. 2 The threshold of SPE. Among them, the statistical indicator T... 2 The SPE (Statistical Parameters) represents the distance of the projection of the variable data in the principal component space from the model center of the PCA normal fluctuation model. The SPE represents the projection of the variable data in the vertical space (noise space) of the principal component space of the PCA normal fluctuation model, that is, the distance from the PCA normal fluctuation model space.
[0080] After data cleaning, the corresponding PCA algorithm is run for offline modeling to describe the linear relationship between the fluctuation values of various variables under normal fluctuation conditions. The specific steps include:
[0081] S21, Calculate the covariance matrix of matrix X. In this embodiment, the covariance matrix of the matrix is calculated using the following formula:
[0082] ;
[0083] In the formula, matrix C is the covariance matrix, and N is the number of time points in matrix X.
[0084] S22, Perform eigenvalue decomposition on the covariance matrix and arrange the eigenvalues in descending order. In this embodiment, the eigenvalue decomposition and arrangement of the covariance matrix are performed according to the following formula:
[0085] ;
[0086] In the formula, , These are the eigenvalues.
[0087] The purpose of eigenvalue decomposition is to use eigenvalues To describe how each principal component can capture information from the original dataset, after sorting the feature values in descending order... > ,…> m is the number of columns in matrix X, which is the number of variable fluctuation values; the characteristic matrix P consists of eigenvalues. The corresponding feature vector composition:
[0088] and , Let be the i-th element in the characteristic matrix.
[0089] S23. Calculate the percentage of information captured by each principal component based on the eigenvalues, and select the number of principal components based on this to determine the loading matrix, which is the core of the PCA model.
[0090] Due to eigenvalues This describes the information that each principal component can capture from the variable dataset; therefore, the percentage of information that each principal component can capture is calculated using the following formula:
[0091] ;
[0092] In the formula, PC i Let i be the i-th principal element.
[0093] because The data has already been sorted in descending order, so the percentage of total information captured by the first k principal elements is:
[0094] ;
[0095] The method for selecting the number of principal components is as follows: If the total information captured by the first k principal components is... Th is a predetermined information threshold, representing the boundary between information and noise; then k is used as the number of principal components. For example... Figure 3 As shown, PCA calculations were performed on the fluctuation values of 11 variables. Th was chosen to be 95. When the number of principal components was 7, the percentage of the total information captured exceeded 95%. Therefore, the number of principal components could be chosen as k=7.
[0096] Once the number of principal components is determined, the loading matrix V of the principal component space can be determined based on the feature matrix P. The loading matrix V is the first k columns of data from the feature matrix P. The loading matrix is the core of the PCA model; it is the transformation matrix that projects variable data into the principal component space. Subsequently, the corresponding statistical indicator T can be calculated based on the loading matrix. 2 , SPE, and the corresponding threshold.
[0097] S24, Calculate statistical index T based on the loading matrix. 2 The threshold corresponding to SPE.
[0098] from Figure 4 As can be seen from this, the statistical indicator T 2The variable T represents the distance from the projection of the variable data into the principal component space to the model center. The SPE index represents the projection of the variable data into the perpendicular space (noise space) of the principal component space, i.e., the distance from the principal component model space. Therefore, the index T can be calculated. 2 The threshold of SPE is used to determine T. 2 Warnings may be issued based on data indicating that SPE exceeds limits.
[0099] In this embodiment, the statistical indicator T 2 threshold The calculation method is as follows:
[0100] ;
[0101] In the formula, m is the number of columns in matrix X, representing the number of variables in matrix X; It follows an F-distribution with the first degree of freedom k and the second degree of freedom mk, therefore Based on confidence level A function for calculating the confidence space;
[0102] Threshold of the statistical indicator SPE The calculation method is as follows:
[0103] ;
[0104] ;
[0105] In the formula, Let be the quantiles on the standard normal distribution corresponding to the desired confidence level α, and m be the number of columns in matrix X.
[0106] Thus far, through steps S21-S24, the PCA normal fluctuation model for the fluctuation values of various variables in the wastewater treatment process has been established. The established PCA normal fluctuation model will be applied online to monitor the fluctuations in the wastewater treatment process in real time, specifically including the following steps:
[0107] S3 collects various variables related to the operation of the reaction process during wastewater treatment, serving as real-time variable data.
[0108] In this embodiment, the variable types of real-time variable data and historical variable data are consistent, forming a vector of real-time variable data at a certain moment. Subsequently, the real-time variable data is cleaned according to the data cleaning method described in step S1.
[0109] S4. Input the real-time variable data into the established PCA normal fluctuation model and calculate the statistical index T. 2 SPE, and the calculated statistical index T 2 SPE and statistical indicator T2 The threshold of SPE is compared, and if the calculated statistical index T is... 2 If the SPE value exceeds the corresponding threshold, an early warning will be issued.
[0110] In this embodiment, the cleaned real-time variable data is input into the established PCA normal fluctuation model to calculate the statistical index T. 2 SPE, and the calculated statistical index T 2 SPE and statistical indicator T 2 The thresholds for SPE are compared. Specifically, the statistical indicator T... 2 The calculation method is as follows:
[0111] Represent the current real-time variable data x as a row vector, and calculate the score of the current real-time variable data x for each principal component. Then, based on the score, calculate the corresponding real-time variable data x. As shown in the following formula:
[0112] ;
[0113] ;
[0114] Calculate the corresponding real-time variable data x Statistical indicator T determined offline 2 threshold The system performs comparisons and issues warnings for real-time variable data that exceeds limits.
[0115] The statistical indicator SPE is calculated as follows:
[0116] Calculate the error of reconstructing the current real-time variable data x. That is, the information difference between the real-time variable data x after reconstructing it through the principal component space and the real-time variable data x, and the SPE is calculated based on this, as shown in the following formula:
[0117] ;
[0118] .
[0119] Calculate the real-time variable data x corresponding to Statistical indicators determined offline threshold The system performs comparisons and issues warnings for real-time variable data that exceeds limits.
[0120] In this embodiment, when the calculated statistical index T 2 When SPE is greater than the corresponding threshold, the algorithm also calculates the relationship between each variable in the real-time variable data x and T.2 The contribution of SPE is used to determine the source of volatility, thus completing the source tracing of volatility. Among these, the variable pairs with the statistical indicator T... 2 The contribution is calculated as follows:
[0121] The j-th variable in the real-time variable data vector x For the i-th score Contribution for:
[0122] ;
[0123] in To load the elements of matrix V at (i, j);
[0124] Then, by summing, the j-th variable is calculated. For T 2 Contribution for:
[0125] ;
[0126] Contribution of variables to the statistical indicator SPE The calculation method is as follows:
[0127] ;
[0128] in Let be the absolute value of the error e.
[0129] Based on the calculated real-time variable data x, each variable pair T 2 The contribution of SPE, plotting the T data for each variable. 2 The contribution plot of SPE (Special Purpose Parameters) allows for a direct observation of the variables causing the fluctuation data to exceed limits. For example, Figure 5 Display T of 11 variables at a certain moment. 2 Contribution graph, such as Figure 5 As shown, the 10th variable (FLUC(O2,h)) affects T. 2 The indicator with the largest contribution can be explained by the hourly fluctuations in aeration volume, which have the greatest impact on the fluctuations in the current wastewater treatment process.
[0130] This invention addresses the common fluctuations and shocks affecting the stable operation of wastewater treatment processes by proposing a wastewater treatment process fluctuation monitoring method based on principal component analysis (PCA). This method utilizes various variables reflecting the operational status of the wastewater treatment process to establish an offline PCA normal fluctuation model, which is then deployed online to monitor fluctuations in these variables. When fluctuations exceed acceptable limits, an early warning is issued, enabling effective monitoring of abnormal fluctuations in the wastewater treatment process.
[0131] The aforementioned wastewater treatment process fluctuation monitoring method based on principal component analysis, when fluctuations exceed the standard, also calculates the relationship between each variable and T in the real-time variable data. 2 The contribution of SPE is used to determine the source of fluctuations, complete the fluctuation source tracing, find the causal variables that cause the fluctuation data to exceed the limit, and guide operators to deal with fluctuations normally, thereby reducing the further impact of fluctuations and shocks and achieving the goal of stabilizing the wastewater treatment process.
[0132] The above description is a detailed description of the preferred embodiments of the present invention. However, the embodiments are not intended to limit the scope of the patent application of the present invention. All equivalent changes or modifications made under the technical spirit of the present invention should fall within the patent scope covered by the present invention.
Claims
1. A method for monitoring fluctuations in wastewater treatment processes based on principal component analysis, characterized in that, Includes the following steps: S1, collect various variables of the reaction process operation during historical wastewater treatment processes as historical variable data; S2, Based on the collected historical variable data, run the PCA algorithm for offline modeling to establish a PCA normal fluctuation model, and calculate the statistical index T during the establishment of the PCA normal fluctuation model. 2 The threshold of SPE, the statistical indicator T 2 The SPE (Statistical Parameters) represents the distance between the projection of the variable data in the principal component space and the model center of the PCA normal fluctuation model. S3 collects various variables related to the operation of the reaction process during wastewater treatment as real-time variable data; S4, input the real-time variable data into the established PCA normal fluctuation model, and calculate the statistical index T. 2 SPE, and the calculated statistical index T 2 SPE and statistical indicator T 2 The threshold of SPE is compared, and if the calculated statistical index T is... 2 If the SPE value exceeds the corresponding threshold, an early warning will be issued.
2. The wastewater treatment process fluctuation monitoring method based on principal component analysis as described in claim 1, characterized in that, In steps S1 and S3, data cleaning is also performed on the collected historical and real-time variable data. The data cleaning method is as follows: the fluctuation of the variable data in the sewage treatment process is quantified at different time scales to obtain the fluctuation quantification value. After completing the quantification of the fluctuations in the variable data, the time period of normal fluctuation is selected based on the quantified fluctuation values of the variable data. After selecting the time period for normal fluctuations, the fluctuation quantization values of variable data under different sampling frequencies are timestamped and matched to construct matrix X. The fluctuation quantization values are then normalized to form a standard probability distribution with a mean of 0 and a variance of 1.
3. The wastewater treatment process fluctuation monitoring method based on principal component analysis as described in claim 2, characterized in that, The method for quantifying the fluctuation of variable data in the wastewater treatment process across different time scales is as follows: divide the standard deviation of the variable data over a period of time by the mean over that period to obtain the quantified fluctuation value.
4. The wastewater treatment process fluctuation monitoring method based on principal component analysis as described in claim 2, characterized in that, The steps for establishing the PCA normal fluctuation model include: S21, Calculate the covariance matrix of the matrix X; S22, perform eigenvalue decomposition on the covariance matrix and arrange the eigenvalues in descending order; S23, calculate the percentage of information captured by each principal component based on the eigenvalues, and select the number of principal components based on this to determine the loading matrix, which is the core of the PCA model; S24, Calculate statistical index T based on the loading matrix. 2 The threshold corresponding to SPE.
5. The wastewater treatment process fluctuation monitoring method based on principal component analysis as described in claim 4, characterized in that, In step S21, the covariance matrix of the matrix is calculated using the following formula; ; In the formula, matrix C is the covariance matrix, and N is the number of time points in matrix X.
6. The wastewater treatment process fluctuation monitoring method based on principal component analysis as described in claim 4, characterized in that, In step S22, the covariance matrix is decomposed into eigenvalues according to the following formula: ; In the formula, , For eigenvalues; m is the number of columns in matrix X, which is the number of variable fluctuation values; The characteristic matrix P consists of eigenvalues The corresponding feature vector composition: and , Let be the i-th element in the characteristic matrix.
7. The wastewater treatment process fluctuation monitoring method based on principal component analysis as described in claim 4, characterized in that, In step S23, due to the eigenvalues This describes the information that each principal component can capture from the variable dataset; therefore, the percentage of information that each principal component can capture is calculated using the following formula: ; In the formula, PC i Let i be the i-th principal element; because The data has already been sorted in descending order, so the percentage of total information captured by the first k principal elements is: ; The method for selecting the number of principal components is as follows: If the total information captured by the first k principal components is... Th is a predetermined information threshold, representing the boundary between information and noise; then k is taken as the number of principal components; After determining the number of principal components, the loading matrix V of the principal component space can be determined based on the feature matrix P. The loading matrix V is the first k columns of data from the feature matrix P. .
8. The wastewater treatment process fluctuation monitoring method based on principal component analysis as described in claim 4, characterized in that, In step S24, the statistical indicator T 2 threshold The calculation method is as follows: ; In the formula, m is the number of columns in matrix X, representing the number of variables in matrix X; It follows an F-distribution with the first degree of freedom being k and the second degree of freedom being mk. Therefore Based on confidence level A function for calculating the confidence space; Threshold of the statistical indicator SPE The calculation method is as follows: ; ; In the formula, Let be the quantiles on the standard normal distribution corresponding to the desired confidence level α, and m be the number of columns in matrix X.
9. The wastewater treatment process fluctuation monitoring method based on principal component analysis as described in claim 1, characterized in that, In step S4, the statistical indicator T 2 The calculation method is as follows: Represent the current real-time variable data x as a row vector, and calculate the score of the current real-time variable data x for each principal component. Then, based on the score, calculate the corresponding real-time variable data x. As shown in the following formula: ; ; The statistical indicator SPE is calculated as follows: Calculate the error of reconstructing the current real-time variable data x. That is, the information difference between the real-time variable data x after reconstructing it through the principal component space and the real-time variable data x, and the SPE is calculated based on this, as shown in the following formula: ; 。 10. The wastewater treatment process fluctuation monitoring method based on principal component analysis as described in claim 1, characterized in that, In step S4, when the calculated statistical index T... 2 When SPE is greater than the corresponding threshold, the algorithm also calculates the relationship between each variable in the real-time variable data x and T. 2 The contribution of SPE is used to determine the source of volatility, thus completing the volatility origin tracing: Variables versus statistical indicators T 2 The contribution is calculated as follows: The j-th variable in the real-time variable data vector x For the i-th score Contribution for: ; In the formula To load the elements of matrix V at (i, j); Then, by summing, the j-th variable is calculated. For T 2 Contribution for: ; Contribution of variables to the statistical indicator SPE The calculation method is as follows: ; In the formula Let be the absolute value of the error e.