Abnormal water quality early warning method, system and device based on constant false alarm and medium

By adaptively determining the number K of Gaussian mixture distributions, and combining the Neyman-Pearson criterion and the bisection method to calculate and dynamically adjust the early warning threshold, the problems of high false alarm rate and untimely warning in traditional water quality early warning methods are solved, thus achieving timeliness and accuracy of water quality early warning.

CN121034037APending Publication Date: 2025-11-28PENGXI SEMICONDUCTOR TECHNOLOGY (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511225334.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Traditional methods for early warning of abnormal water quality often rely on fixed threshold settings that are difficult to adapt to different production environments, leading to high false alarm rates or untimely warnings, which in turn affect equipment stability and pose environmental risks.

Method used

A constant false alarm rate (CFAR) method is adopted, which adaptively determines the number K of Gaussian mixture distributions, and combines the Neyman-Pearson criterion and the bisection method to dynamically adjust the warning threshold, thereby achieving a warning with the false alarm rate controlled within an acceptable range.

Benefits of technology

It enables dynamic adjustment of thresholds based on water quality changes, reduces false alarm rates, provides timely warnings of abnormal water quality, and ensures equipment stability and environmental efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121034037A_ABST
    Figure CN121034037A_ABST
Patent Text Reader

Abstract

The invention discloses an abnormal water quality early warning method, system and device based on constant false alarm and a medium, and the method comprises the steps: obtaining historical data of water quality indexes, and constructing a real sequence of the water quality indexes; determining the number K of Gaussian distributions in the Gaussian mixture distribution of the real sequences through an adaptive method, wherein the probability density function of any real sequence is represented by the sum of K Gaussian distributions; calculating a probability density function of a real sequence in Gaussian mixture distribution through a Neyman-Pearson criterion and a dichotomy to obtain a threshold value of an upper limit and a threshold value of a lower limit of judgment of a predetermined false alarm rate; and performing early warning on the water quality according to standards of an upper limit threshold value and a lower limit threshold value. The probability density function of the water quality index is described through Gaussian mixture distribution. The number K of Gaussian distribution in Gaussian mixture distribution is determined through a self-adaptive method, the accuracy of the K value is measured in combination with the statistical magnitude H and the statistical magnitude U, and the threshold calculation process is simplified through a dichotomy. And the false alarm rate is controlled within an acceptable range.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of abnormal water quality early warning, in particular to an abnormal water quality early warning method, system and device based on constant false alarm and a medium. BACKGROUND

[0002] In chemical, power, steel, papermaking, textile and semiconductor industries, the quality of water directly affects the stability of production process, the service life of equipment and the efficiency of energy use. Impurities, corrosive substances and sediments in water can cause equipment damage and low efficiency. At the same time, many industrial waters are recycled, if the water quality is not monitored and adjusted in time, it may affect the efficiency of water reuse, and even bring environmental risks.

[0003] Traditional abnormal water quality early warning usually compares the sampling results with a fixed threshold, which has many problems. If the threshold is set too small, the false positive rate will be high; if the threshold is set larger, the water quality may not meet the requirements by the time of early warning, and even some precision instruments may have been irreversibly damaged. In addition, the fixed threshold is usually set according to experience, which is difficult to adapt to the needs of different production environments. SUMMARY

[0004] In view of the above technical problems, the purpose of the present application is to provide an abnormal water quality early warning method based on constant false alarm, which can dynamically adjust the early warning threshold to control the false positive rate within an acceptable range of the system, and timely early warn abnormal water quality.

[0005] The abnormal water quality early warning method based on constant false alarm of the present application comprises: S1, obtaining historical data of water quality indicators to construct real sequences of water quality indicators; S2, determining the number K of Gaussian distributions in the Gaussian mixture distribution of the real sequence by an adaptive method, wherein the probability density function of any real sequence is represented by the sum of K Gaussian distributions; S3, calculating the threshold value of the upper limit and the threshold value of the lower limit of the decision of the predetermined false alarm rate by the Neyman-Pearson criterion and the bisection method to calculate the probability density function of the real sequence in the Gaussian mixture distribution; S4, using the upper limit threshold and the lower limit threshold as the standard to early warn the water quality.

[0006] Step S2 comprises: S2.1, calculating the parameters of the Gaussian mixture distribution by the expectation maximization algorithm; S2.2, randomly generating a random sequence of the Gaussian mixture distribution according to the parameters; S2.3, calculating the H statistic and the U statistic of the generated random sequence and the real sequence and , △ and △ And obtain the optimal K value. and the optimal K value ; in, The H-statistic is the Kruskal-Wallis test for the random sequence and the true sequence when K=k; The U-statistic for the Mann-Whitney test of random and true sequences when K=k; The optimal K value under the Kruskal-Wallis test results; This is the optimal K value under the Mann-Whitney test.

[0007] S2.4, Optimize based on the length of the true sequence and best Perform fusion and obtain the K value in the Gaussian mixture distribution.

[0008] Step S2.1 includes: S2.1.1, Initialize the initial parameters of the real sequence ; S2.1.2, based on the current parameters Calculate the expected value of the latent variables = ; Where θ is the parameter to be estimated, X is the data of the observed random variable, and Z is the data of the latent random variable. The log-likelihood function for complete data; S2.1.3, Maximizing the expected log-likelihood function To obtain new parameters = ; S2.1.4, if If the change is small enough or the preset maximum number of iterations is reached, the algorithm terminates; otherwise, it returns to S2.1.2 to continue iterating until... The calculation ends when the change is small enough or the preset maximum number of iterations is reached.

[0009] Step S2.3 includes: S2.3.1 combines the real data and the random data generated by the estimated Gaussian mixture distribution, and ranks them in ascending order; S2.3.2, calculate the sum of the rankings of the actual data, the expression is:

[0010] in, This represents the rank of the j-th data point in the real data sample. Indicates the sample length of the actual data; S2.3.3, calculate the sum of the rankings of the random data generated by the estimated Gaussian mixture distribution, expressed as:

[0011] in, This represents the rank of the j-th data point in a random data sample generated by a Gaussian mixture distribution. This represents the sample length of random data generated by a Gaussian mixture distribution. S2.3.4, Calculate the H statistic, the expression is:

[0012] Where N= + ; If △ Two consecutive times less than 0.1 times the value of K is used to obtain the optimal K value under the H statistic. The H statistic will not be calculated in subsequent iterations; S2.3.5, Calculate the U statistic, which is expressed as follows:

[0013]

[0014]

[0015] If △ Two consecutive times less than 0.1 times the value of K is used to obtain the optimal K value under the U statistic. Subsequent iterations will not calculate the U statistic.

[0016] In step S2.4, the optimal sequence is determined based on the length of the actual sequence. and best To perform the fusion, the expression is:

[0017] Where L is the length of the actual data. for and The minimum value in, for and The maximum value in the Gaussian mixture distribution is used; the value of K after fusion is rounded down to the nearest integer and used as the K value in the Gaussian mixture distribution.

[0018] Step S3 includes: S3.1, calculate the threshold for each of the K Gaussian distributions according to the predetermined false alarm rate; S3.2, Substitute the calculated K thresholds into the Gaussian mixture distribution model to calculate the K false alarm rates; S3.3, sort the K actual false alarm rates and the predetermined false alarm rates together; S3.4 Select the threshold corresponding to one value before and one value after the predetermined false alarm rate; S3.5, continuously update the threshold using the binary search method until the actual false alarm rate equals the predetermined false alarm rate.

[0019] In step S3.1, if the indicator exceeds the upper limit threshold... For the warning, the threshold calculation formula for the Gaussian distribution is:

[0020] If the indicator is below the lower limit threshold For the warning, the threshold calculation formula for the Gaussian distribution is:

[0021] If both the upper and lower limits are set as warnings simultaneously, the threshold calculation formula for the Gaussian distribution is:

[0022]

[0023] in, It is the mean of a Gaussian distribution. It is the variance of the Gaussian distribution. Given a false alarm rate, It is the inverse function of the cumulative distribution function of the Gaussian distribution. ; It is the inverse function of the error function, and the expression of the error function is:

[0024] Given false alarm rate Threshold with upper limit and the lower limit threshold satisfy:

[0025]

[0026] in, It is the probability density function corresponding to the indicator. satisfy:

[0027] in, The weights are the weights of the k-th Gaussian distribution, satisfying 0 ≤ ≤1 and , It is the probability density function of the k-th Gaussian distribution, expressed as:

[0028] In the formula, It is the mean vector of the k-th Gaussian distribution. is the covariance matrix of the k-th Gaussian distribution, and d is the dimension of the data.

[0029] An abnormal water quality early warning system based on constant false alarm rate includes: The data acquisition module is used to obtain historical data on water quality indicators and construct the real sequence of water quality indicators. The determination module is used to determine the number K of Gaussian distributions in the Gaussian mixture distribution of the true sequence through an adaptive method, where the probability density function of any true sequence is represented by the sum of K Gaussian distributions; The calculation module is used to calculate the threshold for the decision using the Neyman-Pearson criterion and the dichotomy method; The early warning module is used to issue early warnings about water quality based on thresholds.

[0030] An abnormal water quality early warning device based on constant false alarms (CFAM) includes: a memory storing a program for an abnormal water quality early warning method based on CFAM and a processor for running the program for an abnormal water quality early warning method based on CFAM, wherein the program for an abnormal water quality early warning method based on CFAM is configured to implement the steps of an abnormal water quality early warning method based on CFAM.

[0031] A computer-readable storage medium storing a program for an abnormal water quality early warning method based on constant false alarm rate (CFAR), wherein when the program is executed by a processor, the steps of the abnormal water quality early warning method based on CFAR are implemented.

[0032] This invention uses a Gaussian mixture distribution to describe the probability density function of water quality indicators. An adaptive method is used to determine the number K of Gaussian distributions in the mixture distribution. The accuracy of K is measured by combining H and U statistics, and a bisection method is used to simplify the integral solution process for threshold calculation. Therefore, this invention solves the problem of relying solely on empirical values ​​for abnormal water quality warnings. Attached Figure Description

[0033] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings.

[0034] Figure 1 This is a flowchart of the abnormal water quality early warning method based on constant false alarm rate in an embodiment of the present invention; Figure 2 This is a flowchart illustrating the steps of the abnormal water quality early warning method based on constant false alarm rate in this embodiment of the invention, which uses an adaptive method to determine the number K of Gaussian distributions in a Gaussian mixture distribution. Figure 3 This is a flowchart illustrating the steps of the abnormal water quality early warning method based on constant false alarm rate in this embodiment of the invention, which uses the Neyman-Pearson criterion and the dichotomy method to solve for the decision threshold. Figure 4 This is a schematic diagram comparing the given false alarm rate obtained from actual data and the actual false alarm rate used in the embodiments of the present invention. Detailed Implementation

[0035] The present invention, based on a constant false alarm rate (CFAR) method for early warning of abnormal water quality, will be further described in detail below with reference to the accompanying drawings. In the following detailed description, only certain exemplary embodiments of the invention are described by way of illustration. Undoubtedly, those skilled in the art will recognize that various modifications can be made to the described embodiments without departing from the spirit and scope of the invention. Therefore, the drawings and description are illustrative in nature and not intended to limit the scope of the claims.

[0036] Time series forecasting is a key technology in fields such as finance, meteorology, and energy. For the aforementioned scenarios, this invention proposes an abnormal water quality early warning method based on constant false alarm rate (CFAR). This method adaptively determines the input features through lag correlation, then performs nonlinear fitting using ensemble learning, and finally filters features again based on their importance.

[0037] The method for issuing abnormal water quality warnings with constant false alarms requires the probability density functions of various water quality indicators, followed by the calculation of thresholds using the Neyman-Pearson criterion. To address this scenario, this invention proposes an adaptive method for estimating the probability density function of the number of Gaussian distributions, combined with the Neyman-Pearson criterion for abnormal water quality warnings. The main steps include: ① Determining the number of Gaussian distributions using an adaptive method; ② Using the H-statistic of the Kruskal-Walls test and the U-statistic of the Mann-Whitney test as a measure between the true distribution and the estimated Gaussian mixture distribution of each indicator; ③ Calculating the threshold using the Neyman-Pearson criterion and the dichotomy method, ensuring that the probability of a false alarm remains constant during the warning process.

[0038] ① Determine the number of Gaussian distributions using an adaptive method. The probability density function of water quality indicators was fitted using a Gaussian mixture distribution. a. The Gaussian mixture distribution used is a mixture of K Gaussian distributions, as shown below:

[0039] in, The weights are the weights of the k-th Gaussian distribution, satisfying 0 ≤ ≤1 and , It is the probability density function of the k-th Gaussian distribution, expressed as:

[0040] In the formula, It is the mean vector of the k-th Gaussian distribution. Let be the covariance matrix of the k-th Gaussian distribution, and d be the dimension of the data. If we are calculating the joint distribution of n indicators, then d = n; if we are calculating the distribution of a single indicator, then d = 1.

[0041] b. Estimate the Gaussian mixture distribution for different K values ​​using the EM (Expectation-maximization algorithm), then generate random sequences according to the estimated probability density function, and calculate the Kruskal-Walls test results and Wasserstein distance between the random sequences and the true sequences corresponding to each K value.

[0042] cK starts from 1 and increases by 1 each time, but cannot exceed the length of the actual data.

[0043] d. Starting with K=2, it is necessary to calculate the difference between the current H statistic of the Kruskal-Walls test and the U statistic of the Mann-Whitney test and the previous one;

[0044]

[0045] In the formula, The H-statistic represents the Kruskal-Walls test for the random sequence and the true sequence when K=k. U-statistic represents the Mann-Whitney test for random and true sequences when K=k.

[0046] e. To reduce computational complexity, early stopping is implemented, with the following rules: If twice in a row Less than The loop ends when the value of k is 0.1 times the value of Kruskal-Walls, and k is considered to be the optimal K value under the Kruskal-Walls test results, denoted as K. Similarly, if twice in a row Less than The loop ends when the value is 0.1 times the value of k, and k is considered to be the optimal K value under the Mann-Whitney test, denoted as k. .

[0047] f. Based on the length of the actual data and The fusion process is as follows:

[0048] In the formula, L is the length of the actual data. for and The minimum value in, for and The maximum value in.

[0049] The fused K is rounded down to obtain the K value in the Gaussian mixture distribution.

[0050] ② The H statistic of the Kruskal-Walls test and the U statistic of the Mann-Whitney test are used as measures between the true distribution and the estimated Gaussian mixture distribution of each index.

[0051] The H-statistic of the Kruskal-Walls test is used as a measure of the relationship between the true distribution and the estimated Gaussian mixture distribution of each indicator. The steps are as follows: a. Combine the actual data and the random data generated by the estimated Gaussian mixture distribution, and rank them in ascending order. If there are identical data points, use the average ranking, which is the mean of the rankings of all identical data points.

[0052] b. Calculate the sum of the rankings of the actual data, as shown below:

[0053] in, This represents the rank of the j-th data point in the real data sample. This indicates the sample length of the actual data.

[0054] c. Calculate the sum of the rankings of the random data generated by the estimated Gaussian mixture distribution, as shown below:

[0055] in, This represents the rank of the j-th data point in a random data sample generated by a Gaussian mixture distribution. This represents the sample length of random data generated by a Gaussian mixture distribution. d. Calculate the H statistic, which is expressed as follows:

[0056] in, .

[0057] e. Calculate the U statistic, which is expressed as follows:

[0058]

[0059]

[0060] ③ By calculating the threshold using the Neyman-Pearson criterion and the dichotomy method, the probability of a false alarm during an early warning is kept constant. The threshold is calculated using the Neyman-Pearson criterion and the dichotomy method to ensure that the probability of a false alarm is constant during an early warning. The specific steps are as follows: a. For warnings where indicators exceed the upper limit of the standard value, the expression for the relationship between the false alarm rate and the threshold is:

[0061] in, It is the probability density function corresponding to the indicator. It is the threshold of the upper limit. This is the false alarm rate, ranging from 0 to 1. It can be set according to the system's tolerance for false alarms. A higher tolerance will increase the false alarm rate, up to a maximum of 1; a lower tolerance will decrease the false alarm rate, down to a minimum of 0.

[0062] b. For warnings where the indicator is below the lower limit of the standard value, the expression for the relationship between the false alarm rate and the threshold is:

[0063] in, It is the probability density function corresponding to the indicator. It is the threshold of the lower limit.

[0064] c. Calculate the threshold using the binary search method. and The process is as follows: The threshold for each Gaussian distribution is calculated individually based on the given false alarm rate, with the upper and lower limits calculated separately.

[0065] Substitute these thresholds into a Gaussian mixture distribution to calculate the actual false alarm rate under the Gaussian mixture distribution, and then sort these calculated results together with the given false alarm rate in ascending order.

[0066] Find the threshold corresponding to one value before and one value after the given false alarm rate, and then use the binary search method to iterate until the actual false alarm rate equals the given false alarm rate.

[0067] Example Figure 1 This is a schematic diagram of the overall process of the abnormal water quality early warning method based on constant false alarm rate in an embodiment of the present invention.

[0068] Step S1: Obtain historical data of water quality indicators for each system, remove values ​​outside the standard range, and construct the true sequence of each indicator.

[0069] In this embodiment, the data used for rationality verification was a water quality dataset from a drinking water company. This dataset contains water quality indicators for 3267 different water bodies, of which 1278 samples are suitable for human consumption and 1998 are not. Data examples are shown in Table 1.

[0070] Table 1

[0071] As shown in Table 1, water quality indicators generally include multiple parameters. In this embodiment, pH is selected as the indicator to specifically illustrate an abnormal water quality early warning method based on constant false alarm rate (CFAR). Correspondingly, the issue that needs to be considered is that the pH indicator contains abnormal data with particularly low or high pH values. In order to construct a reasonable sequence, it is necessary to consider removing abnormal data.

[0072] Considering that the WTO recommends that the pH value should be between 6.5 and 8.5, outliers outside this range were removed when constructing the data sequence.

[0073] The anomaly warning problem can be reduced to a binary detection problem, which results in two types of errors: false alarms and false negatives, and these two errors cannot be reduced simultaneously. To address this issue, the Neyman-Pearson criterion was proposed, which minimizes the false alarm rate while maintaining an acceptable false alarm rate.

[0074] Applying this criterion requires knowledge of the probability density function of normal data. The probability density function of any sequence can be represented by the sum of K Gaussian distributions, i.e., a Gaussian mixture distribution. Its expression is:

[0075] in, The weights are the weights of the k-th Gaussian distribution, satisfying 0 ≤ ≤1 and , It is the probability density function of the k-th Gaussian distribution, expressed as:

[0076] In the formula, It is the mean vector of the k-th Gaussian distribution (with length d). Let be the covariance matrix of the k-th Gaussian distribution, and d be the dimension of the data. If we are calculating the joint distribution of n indicators, then d = n; if we are calculating the distribution of a single indicator, then d = 1.

[0077] Step S2: Determine the number K of Gaussian distributions in the Gaussian mixture distribution using an adaptive method.

[0078] This step S2 includes 8 sub-steps, as shown in the flowchart below. Figure 2 As shown.

[0079] Step S21: Use the EM algorithm to estimate the Gaussian mixture distribution for different K values.

[0080] The specific process of the EM algorithm is as follows: First, initialize the initial parameters of the real sequence. This includes the number of Gaussian distributions K, as well as the mean and standard deviation of each Gaussian distribution.

[0081] Then, perform step E, based on the current parameters. Calculate the expected value of the latent variables: = ; Where θ represents the parameters to be estimated, which are the weights, mean, and variance of each Gaussian distribution. If there are n Gaussian distributions, the parameters to be estimated are 3n. X is the observed random variable data (here, historical data), and Z is the latent random variable data (here, which Gaussian distribution a particular historical data point comes from). Log p It is a whole, representing the log-likelihood function of complete data. x and z together represent complete data, while x is also called incomplete data.

[0082] Next, perform M steps to maximize the expected log-likelihood function. The new parameters are obtained: = ; Finally, if If the change is small enough or the preset maximum number of iterations is reached, the algorithm terminates; otherwise, it returns to step E and continues iterating until... The calculation ends when the change is small enough or the preset maximum number of iterations is reached.

[0083] Step S22: Randomly generate a Gaussian mixture distribution sequence according to the parameters estimated in step S21.

[0084] Step S23: Calculate the H statistic of the generated random sequence and the real sequence. And calculate its difference from the previous one.

[0085] The specific calculation process for the H statistic is as follows: First, the actual data and the random data generated by the estimated Gaussian mixture distribution are combined and ranked in ascending order. If there are duplicate data, the average ranking is used, which is the mean of the rankings of all identical data.

[0086] Then, the sum of the actual data rankings is calculated, which is represented as follows:

[0087] in, This represents the rank of the j-th data point in the real data sample. This indicates the sample length of the actual data.

[0088] Next, the sum of the rankings of the random data generated by the estimated Gaussian mixture distribution is calculated, as follows:

[0089] in, This represents the rank of the j-th data point in a random data sample generated by a Gaussian mixture distribution. This represents the sample length of random data generated by a Gaussian mixture distribution.

[0090] Finally, the H statistic is calculated, and it is expressed as follows:

[0091] in, .

[0092] Step S24, if Two consecutive times less than 0.1 times the value of K is used to obtain the optimal K value under the H statistic. Subsequent iterations will not execute step S23.

[0093] Step S25: Calculate the U statistic for the generated random sequence and the real sequence. And calculate its difference from the previous one.

[0094] The specific calculation process for the U statistic is as follows: First, the actual data and the random data generated by the estimated Gaussian mixture distribution are combined and ranked in ascending order. If there are duplicate data, the average ranking is used, which is the mean of the rankings of all identical data.

[0095] Then, the sum of the actual data rankings is calculated, which is represented as follows:

[0096] in, This represents the rank of the j-th data point in the real data sample. This indicates the sample length of the actual data.

[0097] Next, the sum of the rankings of the random data generated by the estimated Gaussian mixture distribution is calculated, as follows:

[0098] in, This represents the rank of the j-th data point in a random data sample generated by a Gaussian mixture distribution. This represents the sample length of random data generated by a Gaussian mixture distribution.

[0099] Finally, the U statistic is calculated, and it is expressed as follows:

[0100]

[0101]

[0102] Step S26: If Two consecutive times less than 0.1 times the value of K is used to obtain the optimal K value under the U statistic. Subsequent iterations will not execute step S25.

[0103] Step S27: and Once all are obtained, the traversal ends; otherwise, return to setting K to K+1 and return to step S21.

[0104] Step S28: Adjust according to the length of the actual sequence and The fusion process is as follows:

[0105] In the formula, L is the length of the actual data. for and The minimum value in, for and The maximum value in the Gaussian mixture distribution is obtained by rounding down the fused K value.

[0106] Step S3: Solve for the decision threshold using the Neyman-Pearson criterion and the dichotomy method.

[0107] In this embodiment, both excessively high and excessively low pH values ​​are considered abnormal and require an alert. Therefore, this is a two-way binary test problem, requiring the calculation of two thresholds, one large and one small. An alert is needed when the pH value is below the smaller threshold or above the larger threshold.

[0108] For warnings where indicators exceed the upper limit of the standard value, the expression for the relationship between the false alarm rate and the threshold is as follows:

[0109] in, It is the probability density function corresponding to the indicator. It is the threshold of the upper limit. This is the false alarm rate, ranging from 0 to 1. It can be set according to the system's tolerance for false alarms. A higher tolerance will increase the false alarm rate, up to a maximum of 1; a lower tolerance will decrease the false alarm rate, down to a minimum of 0.

[0110] For warnings where indicators fall below the lower limit of the standard value, the expression for the relationship between the false alarm rate and the threshold is as follows:

[0111] in, It is the probability density function corresponding to the indicator. It is the threshold of the lower limit.

[0112] This step S3 includes 5 sub-steps, as shown in the flowchart below. Figure 3 As shown.

[0113] Step S31: Calculate the threshold for each of the K Gaussian distributions according to the given false alarm rate, as shown below: For warnings where indicators exceed the upper limit of the standard value, the Gaussian distribution threshold calculation formula is as follows:

[0114] For warnings where indicators fall below the lower limit of the standard value, the Gaussian distribution threshold calculation formula is as follows:

[0115] For dual-sided warnings, meaning warnings are issued both when the value exceeds the upper limit of the standard and when it falls below the lower limit of the standard, two thresholds need to be calculated:

[0116]

[0117] in, It is the mean of a Gaussian distribution. It is the variance of the Gaussian distribution. It is a given false alarm rate. It is the inverse function of the cumulative distribution function of the Gaussian distribution. = .

[0118] It is the inverse function of the error function, which is defined as follows:

[0119] The next step is to handle warnings for indicators exceeding the upper limit of the standard value and warnings for indicators falling below the lower limit of the standard value separately.

[0120] Step S32: Substitute the calculated K thresholds into the Gaussian mixture distribution model to calculate the K false alarm rates.

[0121] Step S33: Sort the K false alarm rates together with the given false alarm rate.

[0122] Step S34: Select the threshold corresponding to one value before and one value after the given false alarm rate.

[0123] Step S35: Continuously update the threshold using the binary search method until the actual false alarm rate equals the given false alarm rate.

[0124] In this embodiment, to verify the effectiveness of the abnormal water quality early warning method based on constant false alarm rate of the present invention, an experiment was conducted using actual data. The experiment was run on Python software.

[0125] The water quality dataset from a drinking water company was used for testing: pH values ​​were selected from the dataset. Data with pH values ​​outside the range of 6.5 to 8.5 were first removed, and then the algorithm's steps were executed, yielding the results shown in Table 2 below. Table 2

[0126] Figure 4 This is a schematic diagram comparing the given false alarm rate obtained using actual data with the actual false alarm rate in an embodiment of the present invention.

[0127] Figure 4 In the diagram, the straight line represents the ideal situation where the actual false alarm rate equals the given false alarm rate, while the curve represents the actual test results. The experimental results show that the actual achievable false alarm rate can well meet the given false alarm rate requirement, and in most cases, it can be less than or equal to the given false alarm rate.

[0128] Step S4: Issue an early warning for abnormal water quality based on the threshold.

[0129] In this embodiment of the invention, when the given false alarm rate is 0.05, the lower limit of pH anomaly warning is 6.54 and the upper limit is 8.39, which effectively plays the role of early warning.

[0130] Actual data experiments show that the algorithm proposed in this embodiment can effectively achieve constant false alarm rate for early warning of abnormal water quality.

[0131] The preferred embodiments of the present invention have been described in detail above, but the present invention is not limited to the embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention, and these equivalent modifications or substitutions are all included within the scope of this application.

Claims

1. An abnormal water quality early warning method based on constant false alarm rate, characterized in that, include: S1, Obtain historical data of water quality indicators and construct the true sequence of water quality indicators; S2, the number K of Gaussian distributions in the Gaussian mixture distribution of the real sequence is determined by an adaptive method, where the probability density function of any real sequence is represented by the sum of K Gaussian distributions; S3, calculate the upper and lower thresholds of the predetermined false alarm rate by calculating the probability density function of the true sequence in the Gaussian mixture distribution using the Neyman-Pearson criterion and the dichotomy method; S4 uses upper and lower threshold standards to issue early warnings about water quality.

2. The abnormal water quality early warning method based on constant false alarm rate as described in claim 1, characterized in that, Step S2 includes: S2.1, Calculate the parameters of the Gaussian mixture distribution using the expectation-maximization algorithm; S2.2, Generate a random sequence with a Gaussian mixture distribution based on the parameters; S2.3, Calculate the H-statistic of the generated random sequence and the real sequence. and U statistic , △ and △ And obtain the optimal K value. and the optimal K value ; in, The H-statistic is the Kruskal-Wallis test for the random sequence and the true sequence when K=k; The U-statistic for the Mann-Whitney test of random and true sequences when K=k; The optimal K value under the Kruskal-Wallis test results; The optimal K value under the Mann-Whitney test; S2.4, Optimize based on the length of the true sequence and best Perform fusion and obtain the K value in the Gaussian mixture distribution.

3. The abnormal water quality early warning method based on constant false alarm rate as described in claim 2, characterized in that, Step S2.1 includes: S2.1.1, Initialize the initial parameters of the real sequence ; S2.1.2, based on the current parameters Calculate the expected value of the latent variables = ; Where θ is the parameter to be estimated, X is the data of the observed random variable, and Z is the data of the latent random variable. The log-likelihood function for complete data; S2.1.3, Maximizing the expected log-likelihood function To obtain new parameters = ; S2.1.4, if If the change is small enough or the preset maximum number of iterations is reached, the algorithm terminates; otherwise, it returns to S2.1.2 to continue iterating until... The calculation ends when the change is small enough or the preset maximum number of iterations is reached.

4. The abnormal water quality early warning method based on constant false alarm rate as described in claim 3, characterized in that, Step S2.3 includes: S2.3.1 combines the real data and the random data generated by the estimated Gaussian mixture distribution, and ranks them in ascending order; S2.3.2, calculate the sum of the rankings of the actual data, the expression is: in, This represents the rank of the j-th data point in the real data sample. Indicates the sample length of the actual data; S2.3.3, calculate the sum of the rankings of the random data generated by the estimated Gaussian mixture distribution, expressed as: in, This represents the rank of the j-th data point in a random data sample generated by a Gaussian mixture distribution. This represents the sample length of random data generated by a Gaussian mixture distribution. S2.3.4, Calculate the H statistic, the expression is: Where N= + ; ,like Two consecutive times less than 0.1 times the value of K is used to obtain the optimal K value under the H statistic. The H statistic will not be calculated in subsequent iterations; S2.3.5, Calculate the U statistic, which is expressed as follows: If △ Two consecutive times less than 0.1 times the value of K is used to obtain the optimal K value under the U statistic. Subsequent iterations will not calculate the U statistic.

5. The abnormal water quality early warning method based on constant false alarm rate as described in claim 4, characterized in that, In step S2.4, the optimal sequence is determined based on the length of the actual sequence. and best To perform the fusion, the expression is: Where L is the length of the actual data. for and The minimum value in, for and The maximum value in the Gaussian mixture distribution is used; the value of K after fusion is rounded down to the nearest integer and used as the K value in the Gaussian mixture distribution.

6. The abnormal water quality early warning method based on constant false alarm rate as described in claim 5, characterized in that, Step S3 includes: S3.1, calculate the threshold for each of the K Gaussian distributions according to the predetermined false alarm rate; S3.2, Substitute the calculated K thresholds into the Gaussian mixture distribution model to calculate the K false alarm rates; S3.3, sort the K actual false alarm rates and the predetermined false alarm rates together; S3.4 Select the threshold corresponding to one value before and one value after the predetermined false alarm rate; S3.5, continuously update the threshold using the binary search method until the actual false alarm rate equals the predetermined false alarm rate.

7. The abnormal water quality early warning method based on constant false alarm rate as described in claim 6, characterized in that, In step S3.1, if the indicator exceeds the upper limit threshold... For the warning, the threshold calculation formula for the Gaussian distribution is: If the indicator is below the lower limit threshold For the warning, the threshold calculation formula for the Gaussian distribution is: If both the upper and lower limits are set as warnings simultaneously, the threshold calculation formula for the Gaussian distribution is: in, It is the mean of a Gaussian distribution. It is the variance of the Gaussian distribution. Given a false alarm rate, It is the inverse function of the cumulative distribution function of the Gaussian distribution. ; It is the inverse function of the error function, and the expression of the error function is: Given false alarm rate Threshold with upper limit and the lower limit threshold satisfy: in, It is the probability density function corresponding to the indicator. satisfy: in, The weights are the weights of the k-th Gaussian distribution, satisfying 0 ≤ ≤1 and , It is the probability density function of the k-th Gaussian distribution, expressed as: In the formula, It is the mean vector of the k-th Gaussian distribution. is the covariance matrix of the k-th Gaussian distribution, and d is the dimension of the data.

8. An abnormal water quality early warning system based on constant false alarm rate, characterized in that, include: The data acquisition module is used to obtain historical data on water quality indicators and construct the real sequence of water quality indicators. The determination module is used to determine the number K of Gaussian distributions in the Gaussian mixture distribution of the true sequence through an adaptive method, where the probability density function of any true sequence is represented by the sum of K Gaussian distributions; The calculation module calculates the probability density function of the true sequence in the Gaussian mixture distribution using the Neyman-Pearson criterion and the dichotomy method to obtain the upper and lower thresholds of the predetermined false alarm rate. The early warning module provides water quality warnings based on upper and lower threshold standards.

9. An abnormal water quality early warning device based on constant false alarm rate, characterized in that, include: A memory storing a program for an abnormal water quality early warning method based on constant false alarms (CFAM) and a processor for running the program, wherein the program is configured to implement the steps of the abnormal water quality early warning method based on CFAM as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, A computer-readable storage medium stores a program for an abnormal water quality early warning method based on constant false alarm rate (CFAR). When the program is executed by a processor, it implements the steps of the abnormal water quality early warning method based on CFAR as described in any one of claims 1 to 7.