A method for probabilistic screening of water quality risk factors in reclaimed water wetlands

By using the Bartlett normality test for homogeneity of variance and the sequential calculation of the empirical normal distribution function, a screening model was established, which solved the problem of large and complex water quality monitoring data and enabled probabilistic screening and risk assessment of wetland water quality risk factors.

CN119646058BActive Publication Date: 2025-10-31JILIN INST OF WATER RESOURCES SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411697587.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-26
Publication Date
2025-10-31
Estimated Expiration
2044-11-26

AI Technical Summary

Technical Problem

Water quality monitoring data is characterized by a large amount of information, high repetition, and strong randomness, making it difficult to effectively screen the impact of various random factors on water quality risk factors.

Method used

A screening model was established by employing the Bartlett normality test, sequence calculation of the empirical normal distribution function, and probability calculation. The model was then used to screen wetland water quality risk factors through probability and statistics theory.

Benefits of technology

Clearly identify the likelihood and timing of various water quality indicators to help relevant departments effectively understand the risk factors of wetland water quality and assess potential water quality risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119646058B_ABST
    Figure CN119646058B_ABST
Patent Text Reader

Abstract

This invention relates to a method for probabilistic screening of water quality risk factors in reclaimed water wetlands, belonging to the field of aquatic ecological water quality screening technology. The method involves establishing a screening model calculation file, performing a Bartlett normality test for homogeneity of variance, calculating the parameters of the empirical normal distribution function, and using the sequence obtained from the empirical normal distribution function to calculate the probability of screening water quality risk factors. Its advantages lie in addressing the characteristics of water quality monitoring data—large amounts of information, high repetition, strong randomness, and fixed monitoring data locations—by classifying and analyzing the characteristics of the information, comparing common features, and clarifying the probability and timing of occurrence of various water quality indicators to be screened. This algorithm model can assist relevant departments in effectively utilizing monitoring data to understand the possible occurrence of wetland water quality risk factors, clarify the probability and timing of occurrence of various screened water quality indicators, and assess potential water quality risks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention pertains to water ecological water quality screening technology, and particularly relates to a method for probabilistic screening of water quality risk factors in reclaimed water wetlands. Background Technology

[0002] Because environmental pollution and aquatic ecosystem damage directly impact human survival and development, my country is increasingly emphasizing the protection of aquatic ecosystems. After appropriate treatment, wastewater from various polluted waters flows through constructed wetlands. Aquatic plants absorb nitrogen, phosphorus, and other substances from the water, converting them into their own energy. They also absorb heavy metals and toxic substances, playing a role in water purification. This purification improves water quality indicators, meeting specific usage requirements and achieving reclaimed water quality standards suitable for different applications, making the water usable for beneficial purposes. Consequently, reclaimed water quality monitoring and evaluation, as a branch of environmental science, has emerged and developed.

[0003] Due to the massive amount of information and high degree of randomness in water quality monitoring data, which increases dramatically with each year and also involves significant duplication, screening for the impact of various random factors on water quality indicators becomes increasingly difficult. Therefore, it is essential to find a method that can simply and clearly provide conclusions on the impact of various random factors on water quality risk factors using existing monitoring data. Summary of the Invention

[0004] This invention provides a method for probabilistic screening of water quality risk factors in reclaimed water wetlands.

[0005] The technical solution adopted by this invention includes the following steps:

[0006] (1) Establish the file for screening model calculation

[0007] 1) Establish the screening model calculation data file

[0008] I) Represent the monitoring sample data in a matrix with the number of monitoring samples as rows and the number of detection indicators of the samples as columns. There are n water quality monitoring samples, which is called the sample size. Each sample has m monitoring indicators, which is the dimension of the random variable.

[0009] II) To supplement the monitoring sample matrix;

[0010] III) Fill in the calculation data file;

[0011] 2) Sample name file

[0012] Paste the date column into the list of monitored sample data, and rewrite the dates in the first row with the standard values;

[0013] 3) Monitor element name files;

[0014] Arrange the monitoring indicator names in columns;

[0015] (2) Perform the Bartlett test for homogeneity of variance of the normal distribution.

[0016] (3) Calculate the parameters of the empirical normal distribution function for the sequence.

[0017] (4) Probability calculation is performed using a sequence that has been calculated based on the empirical normal distribution function.

[0018] 1) The model convergence condition is that the relative error is less than 50% for probabilities less than 10%;

[0019] 2) For probabilities greater than 10% and less than or equal to 25%, the relative error is taken as <30%;

[0020] 3) For probabilities greater than 25% and less than or equal to 50%, the relative error is taken as <20%;

[0021] 4) For probabilities greater than 65%, the relative error should be less than 5%;

[0022] (5) Probability of screening water quality risk factors

[0023] ① The probability that pollution indicators exceed standard values;

[0024] Select the sample probabilities and corresponding monitoring elements for each group of the fifth sequence;

[0025] ② Ranking of pollution indicators exceeding standard values;

[0026] ③ The dates and monitoring indicators for water quality risk factors exceeding the probability threshold β;

[0027] By combining the sample name file, water quality risk factors that exceed the probability threshold can be screened out, along with their probability of occurrence and the corresponding date.

[0028] In step (1) of this invention, the screening model calculation file is established. 1) If the screening model calculation data file is established, then the monitoring sample matrix is ​​represented as an n×m matrix:

[0029] X = (X1, X2, ..., X) j ,…X m ) T

[0030] X j Let j be a column vector, j = 1, 2, ..., m, or

[0031] X = Ax (3)

[0032] A is an n×m constant matrix composed of monitoring data, A = [a1, a2, ..., a...]. m ], aj (j=1,2,…,m) is the column vector representing the monitoring data, x=(x1,x2,…,x m ) T Let be an m-dimensional random variable.

[0033] The steps (1) of this invention involve establishing a screening model calculation file: 1) establishing a screening model calculation data file, and 2) supplementing the monitoring sample matrix with the following content:

[0034] A. Number the samples on the left side of the monitoring data matrix;

[0035] B. Add 3 rows to the top of the matrix: the first row is the name of the monitoring indicator, the second row is the number of the monitoring indicator, and the third row is the standard value of the monitoring indicator.

[0036] C. Add a column to the right side of the matrix to show the monitoring dates of the samples;

[0037] D. Add a row at the bottom of the matrix to define the criteria for values ​​greater than or less than the standard value.

[0038] The steps (1) of this invention involve establishing a file for calculating the screening model: 1) Establishing a data file for calculating the screening model; 2) Filling in the data file.

[0039] The first line of the data file should contain the basic information and calculation requirements of the data, with spaces separating the data. The first data is the number of samples + 1; the second data is the number of monitoring indicators involved in the calculation + 1; the third data is the significance level α value of the Bartlett test; the fourth data is the probability threshold for specifying the probability requirements of the screening risk factors; and the fifth data is the number of probability sequence calculation groups to determine the availability of parameter estimates.

[0040] Step (2) of this invention, performing the Bartlett normality test for homogeneity of variance, includes:

[0041] Suppose that from m normal populations, m samples are randomly drawn independently, and the size of each sample is denoted as n. i The sample variance is The hypothesis test is then:

[0042] The variances of the populations are not all equal.

[0043] Under the condition that H0 is true, the Bartlett test statistic is:

[0044]

[0045] in, This is called the pooled variance;

[0046] Given confidence level 1-α

[0047] If equation (2) is satisfied, then:

[0048] Then H0 is not rejected;

[0049] like Then reject H0 and accept H1;

[0050] If the conclusion is "do not reject satisfying homogeneity of variance", then the test is passed.

[0051] The parameters of the empirical normal distribution function are calculated for the sequence in step (3) of this invention:

[0052] For a sample size of n, for a random variable representing m-dimensional monitoring data, take the number of sequences q = 5, and calculate the expected value of the normal distribution parameters of the sequences using equation (3). and standard deviation

[0053]

[0054] Then calculate the average value and compare the error between the average value and the value obtained by i = n + (q - 1).

[0055] Step (4) of this invention involves calculating the probability using a sequence that calculates the empirical normal distribution function, including:

[0056] Mathematical expectation and standard deviation The value is used as a parameter of the normal distribution, and the data is standardized using the following formula with the standard value as the origin:

[0057]

[0058] Probability calculation using the normal distribution function Φ(x):

[0059]

[0060] In the formula, h is a natural number.

[0061] In step (5) of this invention, the probability of screening water quality risk factors includes the date and monitoring indicators for screening water quality risk factors exceeding the probability threshold β.

[0062] When screening for water quality risk factors, the risk factor threshold β = Φ needs to be determined. -1 The result of (x) can be directly solved using equation (2) for Φ. -1 (x), the method is as follows:

[0063] Let Φ(β) = Φ0

[0064] Construct the constructor g(x) = Φ(x) - Φ0

[0065] Obviously, g(β) = 0

[0066] remember structure

[0067]

[0068] β = Φ can be obtained directly using equation (2). -1 (x);

[0069] Substitute the required probability threshold into equation (2) to obtain the water quality risk factor threshold β. Substitute the water quality risk factor threshold β into equation (16) to obtain the random variable index corresponding to the water quality risk factor threshold β.

[0070]

[0071] That is, satisfying the corresponding condition (16) The monitoring indicator data are the risk factors that exceed the probability threshold, and x ij Standardize by substituting into equation (15), and then substitute into equation (1).

[0072] The advantages of this invention are that, addressing the characteristics of water quality monitoring data—large volume, high repetition, strong randomness, and fixed monitoring data locations—it categorizes and analyzes the features of the information, compares common characteristics, and uses the law of large numbers and the central limit theorem as the sufficient and necessary basis for the algorithm. Applying the uniqueness relationship of characteristic functions, it approximates the distribution function using a sequence empirical distribution function, constructing an algorithm model based on probability and statistics theory that uses the probability of occurrence to screen wetland water quality risk factors. This algorithm can clearly define the probability and timing of occurrence of various water quality indicators to be screened. This algorithm model can assist relevant departments in effectively utilizing monitoring data to understand the possible occurrence of wetland water quality risk factors, clarify the probability and timing of occurrence of various screened water quality indicators, and assess potential water quality risks. Attached Figure Description

[0073] Figure 1 It is a smooth curve that approximates the distribution function using the step curve of the empirical distribution function. Detailed Implementation

[0074] 1. Create a file for screening model calculation.

[0075] I) Establish screening model and calculate data files

[0076] To more clearly illustrate the process of establishing the data matrix for the monitoring samples, a mathematical representation is used. The number of monitoring samples forms the rows, and the number of detection indicators for each sample forms the columns. There are n water quality monitoring samples, which is the sample size. Each sample also has m monitoring indicators, which is the dimension of the random variable. Taking Table 1 as an example, if a certain wetland has 48 monitoring samples and 12 monitoring indicators, meaning the sample size is 48 and the dimension of the random variable is 12, the monitoring data matrix can be represented by a 48*12 matrix.

[0077] Table 1 Ranking of Monitoring Sample Data

[0078]

[0079]

[0080] In probabilistic and statistical terms, the sample size is 48, and the number of monitoring indicators is 12. Table 1 is a 48*12 matrix with column vector X. j j = 1, 2, ..., 12, obviously, X j For an observable random vector, its expected value E(X) j ) and variance D(X j )exist.

[0081] If there are n monitored samples and m monitored indicators, then X j (j = 1, 2, ..., m)

[0082] The monitoring sample matrix can then be represented as an n×m matrix.

[0083] X = (X1, X2, ..., X m ) T

[0084] or

[0085] X = Ax ⑶

[0086] A is an n×m constant matrix composed of monitoring data, A = [a1, a2, ..., a...]. m ], a j (j=1,2,…,m) is the column vector representing the monitoring data, x=(x1,x2,…,x m ) T Let be an m-dimensional random variable.

[0087] II) To make the matrix representation clearer, the monitoring data matrix is ​​represented by a water quality index monitoring data table, and the following information is added to the matrix:

[0088] A. Number the samples on the left side of the monitoring data matrix;

[0089] B. Add 3 rows to the top of the matrix: the first row is the name of the monitoring indicator, the second row is the number of the monitoring indicator, and the third row is the standard value of the monitoring indicator.

[0090] C. Add a column to the right side of the matrix to show the monitoring dates of the samples;

[0091] D. Add a validation row at the bottom of the matrix to validate whether the value is greater than or less than the standard value.

[0092] The monitoring data for calculating the sample detection indicators are shown in Table 3.

[0093] Table 3. Monitoring data of wetland water quality indicators

[0094]

[0095]

[0096] Since the monitoring of reclaimed water wetland water quality is conducted year-round, the monitoring sample data and monitoring time must be accurately labeled when inputting the calculation sample data, and the corresponding calculation content must be included.

[0097] III) Specific steps for filling in the calculation data file:

[0098] Data file names are user-defined. For ease of searching, it is recommended to start with 0. The file name determines the content of the program's calculations. It is recommended to use text files for calculation data, i.e., files with the .txt extension;

[0099] The first line of the data file should contain basic information about the data and calculation requirements. Data can be separated by spaces or commas. The first data point is the number of samples + 1; the second data point is the number of monitoring indicators involved in the calculation + 1; the third data point is the significance level α value of the Bartlett test; the fourth data point is the probability threshold required for screening risk factors; and the fifth data point is the number of probability sequence calculation groups used to determine the availability of parameter estimates.

[0100] The first column is the row number, used to easily view the sample data number; the other columns contain the information specified for the monitoring samples, with each sample occupying one row, and entered in the order specified in the record.

[0101] For example, create a data file by referring to Table 3, and name the file 01.txt;

[0102] According to the input rules for the first line of the data file, the first line of the 01.txt file should contain the following:

[0103] 49,13,0.01,0.75,5

[0104] The first line of input data indicates that there are 49-1=48 rows and 13-1=12 columns. The significance level α of the Bartlett test is 0.01. 0.75 is the probability threshold specified by the screener as 75%, and 5 is the sequence calculation to determine the sample size by selecting 5 sets of sequences to verify whether the probability validity calculation is passed.

[0105] The data below the second line can be simplified: Paste the data from the standard value row of Table 1 to row 48, and from the standard value column to column number 12, into the second line of the 01.txt file, and rewrite "standard value" in the first line with the number "0". This line of data represents the national standard value and is the reference data for whether the water quality meets the standards.

[0106] The last line of the data file needs to be a verification line for the standard values. The line number of the verification line is the same as the first data in the first line. The numbers in the verification line correspond to the data in each column. 0 indicates that the value is less than or equal to the standard value and meets the requirements, and 1 indicates that the value is greater than the standard value and meets the requirements. For example, the national standard requires the value to be greater than a certain value to meet the requirements, so it is represented by 1. The rest are national standard requirements that the value is less than a certain value and meet the requirements, so they are represented by 0.

[0107] The contents of the 01.txt file are shown in Table 4.

[0108] Table 4 01.txt file format

[0109]

[0110]

[0111] 2) Create a sample name file

[0112] The sample name file has a fixed filename that cannot be changed, and its name is: sample name.txt.

[0113] Paste the monitoring date column from Table 3 into the sample name.txt file, and rewrite the monitoring date in the first row with the standard value, because the analysis results will display the calculation results of the corresponding standard value in this position.

[0114] The contents of the sample name.txt file are shown in Table 5.

[0115] Table 5 Sample Names.txt File Format

[0116] Standard value

[0117] 2021-01-07C

[0118] 2021-01-25C

[0119] 2021-03-01C

[0120] 2021-03-16C

[0121] 2021-04-19C

[0122] 2021-05-12C

[0123] 2021-05-28C

[0124] 2021-06-17C

[0125] 2021-07-19C

[0126] 2021-08-10C

[0127] 2021-08-26C

[0128] 2021-09-10C

[0129] 2021-10-22C

[0130] 2021-11-25C

[0131] 2021-12-07C

[0132] 2021-12-23C

[0133] 2022-01-18C

[0134] 2022-02-16C

[0135] 2022-05-10C

[0136] 2022-06-02C

[0137] 2022-06-26C

[0138] 2022-07-04C

[0139] 2022-07-15C

[0140] 2022-08-02C

[0141] 2022-08-21C

[0142] 2022-09-06C

[0143] 2022-09-26C

[0144] 2022-10-24C

[0145] 2022-11-27C

[0146] 2022-12-12C

[0147] 2023-01-08C

[0148] 2023-02-02C

[0149] 2023-03-10C

[0150] 2023-03-20C

[0151] 2023-04-04C

[0152] 2023-04-18C

[0153] 2023-05-18C

[0154] 2023-06-16C

[0155] 2023-07-12C

[0156] 2023-08-11C

[0157] 2023-09-18C

[0158] 2023-10-18C

[0159] 2023-10-27C

[0160] 2023-11-14C

[0161] 2023-11-21C

[0162] 2023-11-28C

[0163] 2023-12-04C

[0164] 2023-12-20C

[0165] 3) Create a file for monitoring element names.

[0166] The file for monitoring element names has a fixed filename that cannot be changed; its name is: element_name.txt.

[0167] Paste the first row of Table 3 into element_name.txt, and delete the sample number and monitoring date fields. Arrange each element in columns, and match the calculation results with the input element names.

[0168] The contents of the element_name.txt file are shown in Table 6.

[0169] Table 6 Element Names.txt file format

[0170] PH

[0171] COD

[0172] ammonia nitrogen

[0173] CadmiumCD

[0174] Total phosphorus

[0175] Total nitrogen

[0176] Dissolved oxygen

[0177] Nitrite nitrogen NO2-N

[0178] Iron (Fe)

[0179] Manganese (Mn)

[0180] Fluoride F ˉ

[0181] Nitrate nitrogen (NO3-N)

[0182] (2) Perform the Bartlett normality test for homogeneity of variance.

[0183] Equation (14) is used to perform a Bartlett test for homogeneity of variance on random variables expressed by m-dimensional monitoring data with a sample size of n.

[0184] The calculation result of this example

[0185] χ 2 =2064

[0186] α = 0.01 degrees of freedom df = 11χ 2 Table value = 14.63

[0187] χ 2 =2064

[0188] Do not reject satisfying homogeneity of variance

[0189] The calculation results show that the Bartlett test is passed.

[0190] (3) Calculate the parameters of the empirical normal distribution function for the sequence.

[0191] For a sample size of n, the number of sequences q = 5 is taken for the random variable expressed by m-dimensional monitoring data;

[0192] In this example, n = 44, m = 12. Considering the case where n = 44, equation (10) represents the capacity.

[0193] i = 44+0, 44+1, ..., 44+4;

[0194] Perform sequence calculations. For each 12-dimensional random variable, calculate the sequence q=5, and then calculate its average value.

[0195] The average value obtained is compared with the value obtained by i = 44 + q - 1 to determine the error.

[0196] The calculation result for this example:

[0197] Table 7. Mathematical Expectation of Normal Distribution Parameters for Sequences and standard deviation

[0198]

[0199] *Note: The column names in this table correspond to the column names in Table 2.

[0200] (4) Probability calculation is performed using the empirical normal distribution function calculated from the sequence.

[0201] Use Table 7 and standard deviation The value is used as a parameter of the normal distribution, with the standard value as the origin, and the data is standardized using the following formula.

[0202]

[0203] Probability calculation using equation (1)

[0204] The calculation result of this example

[0205] Table 8 shows the mathematical expectation calculated from the parameter sequence. and standard deviation Calculate the corresponding probability

[0206]

[0207] *Note: The column names in this table correspond to the column names in Table 2.

[0208] According to the convergence criteria specified in Table 2, all relative errors in Table 8 passed the verification. The computer then displays...

[0209] Based on the principle of the practical impossibility of low-probability events, that is, the idea that low probability has no impact on screening results, a model converges if it meets the following conditions:

[0210] 1. The model convergence condition is that if the probability is less than 10%, the relative error is less than 50%.

[0211] 2. For selections with a probability greater than 10% and less than or equal to 25%, the relative error is <30%.

[0212] 3. For selections with a probability greater than 25% and less than or equal to 50%, the relative error is <20%.

[0213] 4. For probabilities greater than 65%, the relative error should be less than 5%.

[0214] The probabilistic model satisfies the convergence condition!

[0215] If the verification is successful, you can proceed directly to the risk factor screening.

[0216] Note: After clicking "Verify the validity of the probability screening model", steps 3 and 4 will be executed automatically.

[0217] (5) Probability of using the data in this example to screen water quality risk factors

[0218] 1) Start the calculation by inputting the command "Screening of Reclaimed Water Risk Factors". After starting the calculation program, click "Selection and Recognition of Data Files" and input the data file 01.txt as prompted. The file specifies a screening probability threshold of 75%. Substitute the probability threshold of 0.75 into equation (2) to calculate the water quality risk factor threshold β of the probability threshold. Screen the data matrix X of the monitoring index random variable expressed by equation (3) according to the following formula.

[0219]

[0220] Equation (16) shows that the random vector sequence {X} formed by the screening data matrix X j The risk factors for the random vector X, j = 1, 2, ..., m, are controlled by the water quality risk factor threshold β. j In the context of any monitoring indicator x ij (j=1,2,…,n), satisfying The monitoring indicator data represents the risk factors that exceed the probability threshold. In this example, m = 12, n = 48;

[0221] 2) Click on the probability screening of risk factors

[0222] Clicking the "Risk Factor Probability Screening" command in the menu will display the calculated overall probability of water quality risk factors occurring. The calculation results for this example are as follows:

[0223] ① The probability that pollution indicators exceed the standard values

[0224]

[0225] ② Ranking of pollution indicators exceeding standard values

[0226]

[0227] The above calculation results have provided the indicators for contamination exceeding the probability threshold, as shown in Table 9.

[0228] Table 9 Indicators of contamination exceeding probability thresholds

[0229]

[0230] ③ The dates and monitoring indicators for screening water quality risk factors exceeding the probability threshold β.

[0231] Substituting the probability threshold of 0.75 into equation (2) yields the probability threshold β of the water quality risk factor. Substituting the probability threshold β into equation (16) yields the random variable index corresponding to the water quality risk factor threshold β, which is the index that satisfies the condition x in equation (16). ij >β The monitoring indicator data are the risk factors that exceed the probability threshold, and x ij Substituting into equation (15) for standardization, and then into equation (1), combined with Table 5, we can screen out water quality risk factors that exceed the probability threshold, their probability of occurrence, and the corresponding dates. The results are shown in Table 10.

[0232] Table 10. Risk Probability and Date of Possible Pollution Indicators for Sample Elements

[0233]

[0234]

[0235]

[0236] Table 11 Risk Probability and Date of Sample Element Pollution Indicators Exceeding the Probability Threshold

[0237]

[0238]

[0239] Table 12 Results of Probability Screening for Water Quality Risk Factors

[0240]

[0241]

[0242]

[0243] The following section further discusses the construction of the algorithm included in this invention.

[0244] I. Computer Algorithms for Establishing the Forward and Inverse Functions of the Standard Normal Distribution Function

[0245] This invention employs statistical analysis of the parameter sequence calculation of the normal distribution function, and determines the effectiveness of the algorithm by comparing the relative error between the probability of the average parameter and the directly calculated probability. While the calculation process can be performed by looking up tables, it is inefficient; therefore, an algorithm based on the normal distribution function must be established. Furthermore, since the probability of risk factors occurring in the next stage needs to be predicted based on monitoring data, an algorithm based on the inverse normal distribution function must also be established.

[0246] Standard normal distribution function algorithm:

[0247] Using Pade approximation, a confluence superratio function is constructed. An algorithm for establishing the normal distribution function Φ(x) is then developed using a continued fraction recursive scheme.

[0248]

[0249] In the formula, h is a natural number. If h = 16, then the precision of formula (2) can reach more than 6 decimal places, which is sufficient.

[0250] Algorithm for the inverse function of the standard normal distribution function:

[0251] When screening for risk factors, it is necessary to determine the risk factor threshold β = Φ. -1 The inventors used the non-iterative algorithm proposed by Zheng Duo to solve Φ directly using the following formula (x). -1 (x).

[0252] Let Φ(β) = Φ0

[0253] Construct the constructor g(x) = Φ(x) - Φ0

[0254] Obviously, g(β) = 0

[0255] remember structure

[0256]

[0257] This algorithm directly calculates β = Φ without iteration. -1 (x) has a calculation precision of up to 6 decimal places, which is sufficient.

[0258] II. Establish monitoring sample data files

[0259] If we use the number of monitoring samples as rows and the number of detection indicators for each sample as columns, then the monitoring data can be represented by a matrix. To more clearly express the process of establishing the data matrix for the monitoring samples, a mathematical expression is used. There are n water quality monitoring samples, called the sample size, and each sample has m monitoring indicators, which is the dimension of the random variable.

[0260] Its column vector is X j j = 1, 2, ..., 12, obviously, X j For an observable random vector, its expected value E(X) j ) and variance D(X j )exist.

[0261] If there are n monitored samples and m monitored indicators, then X j (j = 1, 2, ..., m)

[0262] The monitoring sample matrix can then be represented as an n×m matrix.

[0263] X = (X1, X2, ..., X m ) T

[0264] or

[0265] X = Ax ⑶

[0266] A is an n×m constant matrix composed of monitoring data, A = [a1, a2, ..., a...]. m ], a j (j=1,2,…,m) is the column vector representing the monitoring data, x=(x1,x2,…,x m ) T Let be an m-dimensional random variable.

[0267] III. Determining the sample size for approximating the empirical distribution function

[0268] If the monitoring data are independent of each other, they are considered as one-dimensional random variables. If the n-dimensional normal distributions are correlated, an orthogonal matrix is ​​established to obtain the eigenvector U, and the overall distribution is decomposed into n independent normal distributions. Thus, the m random variables of the monitoring indicators are considered as independent random variables.

[0269] For simplicity, we assume that the m random variables of the monitoring indicators are mutually independent random variables.

[0270] According to the law of large numbers, if the random variable sequence of reclaimed water wetland water quality monitoring data {X} j If the numbers j = 1, 2, ... are independent and identically normally distributed, then their expected value E(X) is obviously... j Exists, variance 0 <D(X j If ) < +∞, construct the empirical normal distribution function from the samples. Its second moment If ε = 1, 2, ... exists, then for any real number ε > 0, as n → ∞, we have:

[0271]

[0272] This is known as mean square convergence. Equation (4) indicates that as the sample size increases, the second moment of the sample converges to the second moment of the population, meaning that the probability of a difference between the second moment of the sample and the second moment of the population decreases. According to the theorems of probability theory, the expected value of its first moment is also convergent.

[0273] By Glivenko's theorem, the population distribution function of X is F(x), and the empirical distribution function is Fi. n (x), for any real number x, as n→+∞, we have:

[0274]

[0275] Glevenko's theorem states that when the sample size n is sufficiently large, the empirical distribution function F n The probability 1 of F(x) converges to F(x) with respect to x, which is the theoretical basis for inferring the population from the sample. Equation (5) shows that when n is large enough, for all values ​​of x, the empirical distribution function F(x) converges to F(x) with probability 1. n (x) is very close to the population distribution function F(x), such as Figure 1 As shown. This means that, for a random variable, constructing an empirical normal distribution function, when the sample size n is sufficiently large, can be considered that the empirical normal distribution function F... n (x) converges to the normal distribution function F(x), such as Figure 1 At this point, it can be assumed that when the sample size is greater than a certain value n, the empirical distribution function F is such that within the interval where the sample size is greater than n. n The distribution function F(x) is very close to the distribution function F(x), and the parameters of the distribution function are also very close. Based on the uniqueness theorem of characteristic functions, the normal empirical characteristic function is constructed. and the characteristic function of normal distribution The first and second moments should also be very close.

[0276] Figure 1 This intuitively illustrates the relationship between the population's distribution function and its empirical distribution function. Corresponding to equation (5), when n is sufficiently large, for all x values, the empirical distribution function F... n The probability that (x) is very close to the population distribution function F(x) is 1, that is, F n The probability of an event where the absolute value of the difference between F(x) and F(x) is sufficiently small is 1, meaning it will definitely occur. In this case, we can consider that within an interval where the sample size exceeds a certain value n, its empirical distribution function F... n (x) is very close to the distribution function F(x).

[0277] Based on the theorems related to the characteristics of the normal distribution function, when n is sufficiently large, if the empirical characteristic function of q sequences greater than n is... The first and second moments converge to the normal distribution function. If the first and second moments, or in other words, converge to the parameters of the normal distribution function, then the empirical distribution function F of the sequence... n+i The distribution function F(x) uniformly converges to the normal distribution function F(x). This is the theoretical basis for approximating the parameters of the distribution function using an empirical distribution function. Therefore, it is derived that the empirical distribution function F(x)... n The method for determining n when approximating the normal distribution function F(x) is as follows:

[0278] 1. Determine the sample size of the characteristic function parameters of the empirical normal distribution function using the method of moments.

[0279] Since the subvectors of the random vector matrix established from the water quality observation data of reclaimed water wetlands are all independent, obey the law of large numbers, and satisfy the central limit theorem, the method of moments estimates the equations. It uses sample moments to estimate the global moments, but for a single sample...

[0280]

[0281] In the formula, k is the sample order and n is the sample size.

[0282] The method of moments is used to estimate the parameters of the normal distribution. The population follows a normal distribution, and the method of moments equations are established using the k-th moment of the random variable.

[0283]

[0284] The population follows a normal distribution, and its characteristic function is:

[0285]

[0286] Solve for the two parameters of the normal distribution. It is 2-fold differentiable. Based on the characteristics of the characteristic function (8),

[0287]

[0288] Construct the moment method equation using equation (10)

[0289]

[0290] We obtain two parameter estimates for the normal distribution.

[0291]

[0292] because Corrected estimate It is σ 2 The minimum mean square error estimator and the optimal unbiased estimator are typically used in practical applications. As σ 2 The estimate.

[0293] To ensure uniformity of dimensions in calculations, we take...

[0294]

[0295] The present invention uses equation (10) to calculate the parameters of the empirical normal distribution.

[0296] Due to the properties of the characteristic function of the normal distribution function, when n is sufficiently large, if the empirical characteristic function of q sequences greater than n is... The first and second moments converge to the normal distribution function. The first and second moments, then the empirical distribution function F of the sequence. n+i (x) uniformly converges to the normal distribution function F(x). Therefore, we can derive the empirical distribution function F... n An algorithm for approximating the normal distribution function F(x) (x).

[0297] 2. Determining the convergence of the empirical distribution function approximating the distribution function

[0298] From equation (10) To obtain the optimal, unbiased, consistent, efficient, and minimum mean square error estimator, therefore, we adopt... As σ 2 The mean estimate is used to calculate the sample size. Let n be the sample size and q be the number of sets to be calculated. Then the mean estimate is... To ensure the comparability of error estimation results, the dimensions should be the same; therefore, σ needs to be verified. 2 Error is estimated using standard deviation. The distribution function of the nth group is F n Starting with (x), Construct an empirical distribution function for the sequence of parameters (i = n, n+1, n+2, ..., n+q-1). The absolute value of the relative error between the average of the sequence calculation results and the calculation results at i = n + q - 1 is taken as the convergence condition.

[0299]

[0300] Based on the principle of the practical impossibility of low-probability events and the idea that low probability has no impact on screening conclusions, it is considered that if the relative error w satisfies the conditions specified in Table 2 for sequence calculation, then the model is considered to have converged.

[0301] Table 2 Convergence conditions calculated based on probability sequences

[0302] Serial Number The convergence condition for probability sequence calculation (the relative error between the calculated mean and the last term of the sequence) is used. 1 For probability values ​​less than 10%, take w < 50%. 2 For probability values ​​greater than 10% and less than or equal to 25%, take w < 30%. 3 For probability values ​​greater than 25% and less than or equal to 50%, take w < 20%. 4 For probability values ​​greater than 65%, take w < 5%.

[0303] This convergence condition is essentially the estimate of the empirical distribution function parameters of the i=n+q-1th screening indicator in the water quality sample. and The standardized distribution function F is obtained. i The probability value of (x), which is the approximate probability, is the average of the empirical distribution function for q samples with a size greater than or equal to n. The absolute value of the relative error w between them satisfies equation (11).

[0304] Once the set conditions are met, equation (5) is considered valid, meaning the sample size sufficiently meets the conditions of the screening model. This is because the mean and variance of the water quality monitoring data can be directly calculated. Note that... It exists and is greater than 0. It increases as n increases, let it be... If the following formula is true

[0305]

[0306] Then the Lindberg condition, which satisfies the central limit theorem, also satisfies the Lyapunov theorem condition.

[0307] lim n→+∞ P(Y n ≤y)=Φ(y) (13)

[0308] That is, when n is large enough, Y n It converges to a normal distribution with probability. Equations (12) and (13) express that They converge uniformly to 0 in probability, expressing the sum Y. n The terms in the model are uniformly small in probability, and the limiting distribution of the sum of these uniformly small random variables is a normal distribution. These theories constitute the necessary conditions for the probability screening model. When the monitoring sample is taken as a random variable, it satisfies the Lindberg condition (12) when the capacity is large enough. By the Lyapunov theorem of equation (13), Φ(x) can be calculated using equation (1).

[0309] Through extensive calculations, the inventors believe that as long as the convergence conditions specified in Table 2 above are met, the calculation of a sequence with q=5 can satisfy equations (5) and (12) and (13), which means satisfying the sample size of the distribution function parameters for probability screening calculation, and using it to calculate the probability.

[0310] 3. Methods for multidimensional normal distribution

[0311] The probability screening algorithm is based on the law of large numbers and the central limit theorem, which requires that random variables follow a normal distribution. The water quality observation sample of reclaimed water wetland is a set of data composed of multiple monitoring indicators, or in other words, the water quality observation sample is a set of random variables composed of random vectors, with a multivariate function distribution.

[0312] As is well known, the analysis of related multivariate distribution functions is quite complex, and the calculation and solution are difficult, thus limiting its practical application. If it can be proven that independent normal distributions do not affect the overall normal distribution of the sample population, the difficulty of calculation and solution will be significantly reduced, greatly simplifying the problem. Therefore, if it is proven that the independent normality of random vector components does not affect the overall normal distribution, it also proves that the random variation of the water quality monitoring data components in this invention follows the law of large numbers and the central limit theorem, which can be used for probability analysis. This is because only by considering the multiple monitoring indicators of the water quality monitoring sample according to a one-dimensional normal distribution can one conveniently use one-dimensional probability and statistics theory to screen and calculate water quality risk factors.

[0313] The multidimensional sub-vector sequence {X} of the random vector matrix established from the water quality observation data of reclaimed water wetlands j The numbers i = 1, 2, ... are all independent and follow a normal distribution. Their expected value E(X) is... j Existence and variance 0 <D(X j If ) < +∞, by the Lindberg-Levy theorem, the multidimensional subvector sequence {X} j The expression ,j=1,2,…} obeys the central limit theorem.

[0314] This invention applies the uniqueness relation of characteristic functions, that is, the necessary and sufficient condition for two distribution functions F1(x) and F2(x) to be identical is that their characteristic functions are... and The identity theorem proves that an m-dimensional random vector X = (X1, X2, ..., X...) m ) T It consists of monitoring indicators from the observed samples, with each monitoring indicator X j If j = 1, 2, ..., m are mutually independent normal distributions, this does not affect the fact that the sample population X follows a normal distribution, and vice versa. Therefore, for a multidimensional mutually independent normal distribution, it can be decomposed into multiple one-dimensional normal distribution functions to solve.

[0315] If the m-dimensional normal distributions are correlated, the overall distribution can be decomposed into m mutually independent normal distributions by establishing an orthogonal matrix and obtaining the eigenvector U.

[0316] If multiple monitoring indicators of water quality samples are considered according to a one-dimensional normal distribution, then theoretically, the law of large numbers and the central limit theorem can be used to screen water quality risk factors.

[0317] IV. Bartlett's method for testing the homogeneity of variance of a normal distribution

[0318] The Bartlett test is a method used to test whether a sample comes from an uncorrelated population.

[0319] Since the premise of this invention is that the monitoring data follows a normal distribution, it is generally necessary to test the normality and homogeneity of variance of the data before performing analysis of variance, especially when the sample variances differ significantly.

[0320] This invention uses the Bartlett test for variance comparison to test the homogeneity of multiple normal populations (also applicable to two normal populations). Only data that pass the homogeneity test can be used for probability analysis.

[0321] In this section, n represents the size of the monitoring sample, and m represents the dimension of the random variable or the number of populations. The specific method is as follows:

[0322] Suppose that from m normal populations, m samples are randomly drawn independently, and the size of each sample is denoted as n. i The sample variance is Then the hypothesis test is

[0323] The variances of the populations are not all equal.

[0324] Under the condition that H0 is true, the Bartlett test statistic is:

[0325]

[0326] in, This is called pooled variance.

[0327] Given confidence level 1-α

[0328] If equation (14) is satisfied, then

[0329] Then H0 is not rejected;

[0330] like Then reject H0 and accept H1.

[0331] If the conclusion is "do not reject satisfying homogeneity of variance", then the test is passed.

[0332] If the homogeneity test is not passed, it may indicate that a certain test indicator has an outlier. The homogeneity test can be passed by adjusting the sample data.

[0333] After starting the calculation program, click "Bartlett Test" in the menu. If the calculation result displays "Do not reject the satisfaction of homogeneity of variance" on the screen, it indicates that the data has passed the Bartlett test.

[0334] V. Verifying the effectiveness of the probability screening model

[0335] The algorithm is validated for a set of monitoring data. If the probability screening algorithm passes the validation, it means that the algorithm is effective for that set of monitoring data.

Claims

1. A method for probabilistic screening of water quality risk factors in reclaimed water wetlands, characterized in that, Includes the following steps: (1) Establish the file for screening model calculation; 1) Establish the screening model calculation data file; I) Water quality monitoring samples include This is called the sample size, and each sample has one or more monitoring indicators. Let be the dimension of the random variable. The monitoring sample matrix is ​​represented as a matrix with the number of monitored samples as rows and the number of detection indicators per sample as columns. Matrix of order: , … ,or: (3) for A constant matrix of order composed of monitoring data , ( … ) represents the column vector of monitoring data. for 3D random variable; II) To supplement the monitoring sample matrix; III) Fill in the calculation data file; 2) Sample name file; Paste the date column into the list of monitored sample data, and rewrite the dates in the first row with the standard values; 3) Monitor element name files; Arrange the monitoring indicator names in columns; (2) Perform the Bartlett test for homogeneity of variance of the normal distribution; (3) Calculate the parameters of the empirical normal distribution function for the sequence; (4) Probability calculation is performed using a sequence with an empirical normal distribution function: Mathematical expectation and standard deviation The value is used as a parameter of the normal distribution, and the data is standardized using the following formula with the standard value as the origin: (15) Using the normal distribution function Perform probability calculations: (1) In the formula It is a natural number; (5) Probability of screening water quality risk factors: ① The probability that the pollutant index exceeds the standard value; ② Ranking of pollution indicators exceeding standard values; ③ Screening for water quality risk factors exceeding probability thresholds Dates and monitoring indicators; The methods for screening water quality risk factors are as follows: set up = ; Constructor = - ; Obviously, = ; remember = ,structure: (2) We can directly obtain it using equation (2). ; Substitute the required probability threshold into equation (2) to obtain the water quality risk factor threshold of the probability threshold. Water quality risk factor thresholds Substituting into equation (16), the corresponding water quality risk factor thresholds are obtained. Random variable indicators: , , (16) That is, satisfying the corresponding condition (16) The monitoring indicator data are the risk factors that exceed the probability threshold. Standardize by substituting into equation (15), and then substitute into equation (1); By combining the sample name file, water quality risk factors that exceed the probability threshold are screened out, along with their probability of occurrence and the corresponding date.

2. The method for probabilistic screening of water quality risk factors in reclaimed water wetlands according to claim 1, characterized in that: In step (1), the screening model calculation file is established as follows: 1) Establish the screening model calculation data file, 2) Supplement the monitoring sample matrix with the following content: A. Number the samples at the left end of the monitoring data matrix; B. Add 3 rows to the top of the matrix: the first row is the name of the monitoring indicator, the second row is the number of the monitoring indicator, and the third row is the standard value of the monitoring indicator. C. Add a column to the right side of the matrix to represent the monitoring dates of the samples; D. Add a row at the bottom of the matrix to define the criteria for values ​​greater than or less than the standard value.

3. The method for probabilistic screening of water quality risk factors in reclaimed water wetlands according to claim 1, characterized in that: In step (1), the file for establishing the screening model calculation is: 1) Establish the screening model calculation data file; 2) Fill in the calculation data file. The first line of the data file should contain basic information about the data and the calculation requirements. Data should be separated by spaces. The first data point should be the sample size + 1; the second data point should be the number of monitoring indicators used in the calculation + 1; and the third data point should be the significance level determined by the Bartlett test. The fourth data point is the probability threshold required to specify the probability of the screening risk factor; the fifth data point is the number of probability sequence calculation groups used to determine the availability of the parameter estimate.

4. The method for probabilistic screening of water quality risk factors in reclaimed water wetlands according to claim 1, characterized in that: Step (2) of performing the Bartlett normality test for homogeneity of variance includes: Suppose From a normal population, samples are randomly drawn independently. There are n samples, and the size of each sample is denoted as . The sample variance is Then the hypothesis test is: The variances of the populations are not all equal; exist Under the condition that it holds true, the Bartlett test statistic is: (14) in, This is called the pooled variance; Given confidence level ; If equation (2) is satisfied, then: If so, then do not refuse. ; like If so, then refuse. ,accept ; If the conclusion is "do not reject satisfying homogeneity of variance", then the test is passed.

5. The method for probabilistic screening of water quality risk factors in reclaimed water wetlands according to claim 1, characterized in that: The parameters of the empirical normal distribution function are calculated in step (3) of the sequence: With a sample size of ,right The random variable expressed in the monitoring data is the number of sequences. =5, and the expected value of the normal distribution parameter is calculated from equation (3). and standard deviation : (10) Then calculate their average, and the calculated average is compared with... The obtained values ​​are compared for error.

Citation Information

Patent Citations

  • Accelerated degradation data validity testing and model selection method

    CN105468907A

  • Methodology for robust portfolio evaluation and optimization taking account of estimation errors

    US20070288397A1