A method for estimating the average value of pollutant concentration
By replacing data below the detection limit with ω*LOD/2 using a calculated ω, the method addresses the inaccuracy of existing pollutant concentration estimates, achieving accurate and stable pollution level assessments.
Patent Information
- Application Number
- CN202410342862.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-25
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-03-25
AI Technical Summary
Prior art often leads to overestimation or underestimation of the average pollutant level when processing pollutant concentration data below the detection limit, and the average pollutant concentration cannot be accurately estimated.
A method of estimating the mean value of pollutant concentration is used to replace the data with pollutant concentration below LOD with ω*LOD/2. ω is calculated by detecting data with pollutant concentration above LOD. The specific formula of ω is ω=(-229X1-8.29X2-1.44X1X2+1495)/1000, X1=lnA, 0.65≤ω≤1.84.
The accuracy and stability of the estimation of the mean pollutant concentration is improved, and the average concentration of pollutant containing below the detection limit can be simply and accurately estimated.
Smart Images

Figure CN118248251B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of detecting and evaluating pollutant concentrations, and particularly relates to a method for estimating the average value of pollutant concentrations. Background Art
[0002] Exogenous pollutants usually have large concentration variations and high spatial heterogeneity. Therefore, it is often found that the pollutants in some samples in environmental media are below the limit of detection (LOD), and their concentrations cannot be quantified. For example, perfluoroalkyl substances in groundwater, dibenzo[a,h]anthracene in flue gas, benzene in industrial waste land soil, heavy metals in soil, ammonia and phosphorus in water, etc. Environmental observation data containing values below the LOD are usually referred to as "left-censored data". The occurrence of censored environmental observations poses challenges to summarizing statistical characteristics, such as the inability to directly evaluate the average pollution level. To address the challenges posed by censored data, environmental scientists have adopted different statistical strategies to handle this problem.
[0003] For example, the deletion method is adopted: that is, the samples with concentrations below the LOD are deleted and then the average value and standard deviation of the remaining samples are calculated. It has been proven that this is a wrong approach because it will overestimate the average level of pollutants. However, there are still a very small number of studies that calculate the average value after deleting the samples with concentrations below the LOD.
[0004] In addition, there is also the substitution method: that is, the data with concentrations below the LOD are replaced with a specific value and then the average value and standard deviation are calculated. A very small number of studies replace the data with pollutant concentrations below the LOD with 0, which will obviously cause biases. For example, new pollutants of the same type can be divided into multiple monomers. If the concentrations of some monomer pollutants below the LOD are replaced with 0, then the total concentration calculated for these monomers will be on the low side, and the bias cannot be ignored. Therefore, the existing technology needs to be further developed. Summary of the Invention
[0005] In view of the various deficiencies of the prior art, to solve the above problems, a method for estimating the average value of pollutant concentrations is now proposed, and the following technical solutions are provided:
[0006] A method for estimating the average value of pollutant concentrations, the specific steps are as follows: Obtain the detection data of the pollutant concentrations in a plurality of detection samples, and replace the average value of the pollutant concentrations below the LOD with ω*LOD / 2, where ω is obtained from the detection data of the pollutant concentrations above the LOD.
[0007] Further, in the detection samples, the proportion of the detection samples with pollutant concentrations below the LOD is 5%-50%.
[0008] Further, ω has a linear relationship with X1, and X1 = lnA.
[0009]
[0010] Further, ω = (-229X1 - 8.29X2 - 1.44X1X2 + 1495) / 1000, where X2 is the proportional value of the test samples with pollutant concentration lower than the LOD.
[0011] Further, 0.65 ≤ ω ≤ 1.84.
[0012] Beneficial effects:
[0013] 1. The method for estimating the average value of pollutant concentration in the present invention not only has high average accuracy but also is more stable, and is an excellent method for estimating the average value containing values below the detection limit.
[0014] 2. The present invention designs a calculation formula capable of estimating the average value of pollutant concentration below the detection limit by using the known data that can be detected. Using this formula, the average value of pollutant concentration in a certain area can be simply and accurately estimated. Description of the drawings
[0015] Figure 1 It is a diagram showing the influence of the lognormal distribution parameter σlog and <LOD% on the accuracy of the arithmetic mean (am) of the LOD / 2 substitution method in the present invention;
[0016] Figure 2 It is a diagram showing the numerical range of ω in the case of <LOD in the lognormal distribution data in the present invention;
[0017] Figure 3 It is a diagram showing the relationship between ω and the known information (<LOD%, E(X|X≥c), sd(X|X≥c)) in the data containing values below the detection limit in the present invention;
[0018] Figure 4 It is a diagram showing the relationship between ω and lnsd(X|X≥c) / E(X|X≥c) in the present invention;
[0019] Figure 5 It is a diagram showing the accuracy of the average value after replacing the data below the detection limit with LOD / 2 and ω*LOD / 2 proposed by the present invention. Detailed implementation manners
[0020] In order to enable those skilled in the art to better understand the technical solutions of the present invention, the technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Based on the embodiments in this application, other similar embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.
[0021] According to an embodiment of the present invention, a method for estimating the average value of pollutant concentration is provided, and the specific steps are as follows: Obtain the detection data of the pollutant concentration in multiple detection samples, and replace the average value of the pollutant concentration below the LOD with ω*LOD / 2, where ω is obtained from the detection data of the pollutant concentration above the LOD. This method not only has high average accuracy but also is more stable, and is an excellent method for estimating the average value containing values below the detection limit.
[0022] The above estimation method is obtained through the following exploration:
[0023] 1.1 Generation of simulation data
[0024] Use the R language to generate pollutant simulation data that conforms to the log-normal distribution (the true average value and standard deviation are known). By setting the LOD value, convert the values below the LOD into "<LOD" label data, so as to construct simulation data containing values less than the detection limit.
[0025] 1.2 Setting of influencing factors
[0026] Log-normal distribution parameters: The range of ulog variation is [-5, -4, -3, -2, -1, 0, 1, 2, 3, 4, 5],
[0027] The range of σlog variation is [0.2, 0.3, 0.4, 0.6, 0.7, 0.9, 1.2, 1.6, 2.0];
[0028] Sample size (representing the number of investigation samplings): 10, 20, 40, 80, 160 respectively;
[0029] Non-detection rate (<LOD%): By setting the LOD value, the non-detection rate is made 0% to 50%, with an interval of 5%;
[0030] 1.3 Accuracy of the LOD / 2 replacement method
[0031] Using the generated simulation data, study the accuracy of calculating the mean value by the LOD / 2 replacement method under different factors. The formula for calculating the mean accuracy is as follows: Accuracy (%) = mean LOD / 2 / true mean × 100. In this formula, mean LOD / 2 represents the average value calculated after replacing the data of <LOD in the simulation data with LOD / 2, and the true mean is known when generating the data. For pollutant data that conforms to the log-normal distribution, if there are sample pollutant concentrations below the LOD, use the LOD / 2 replacement method to calculate the arithmetic mean value, and the accuracy is between 80% and 101%. As Figure 1 shown, the factors affecting the mean accuracy include σlog and <LOD%, while the distribution parameter ulog and the sample size have no effect on the accuracy geometry, which is one of the discoveries in the exploration process of the present invention. Specifically, as σlog
[0032] With the increase of <LOD%, the accuracy improves. When 0.4 ≤ σlog ≤ 2.0, the accuracy is between 95% and 100%. With the increase of <LOD%, the accuracy rate decreases. That is to say, the more the number of samples with pollutant concentration lower than LOD, the lower the accuracy rate. When the value of <LOD% is less than 10, the accuracy rate reaches more than 95%.
[0033] Numerical characteristics of 1.4ω
[0034] To obtain a more accurate mean value, ω is first defined, and then a weight ω is assigned to LOD / 2 to make the accuracy reach 100%. Through mathematical derivation, it can be obtained that In this formula, is the conditional mean value of <LOD data, which is also the key discovery in the repeated exploration of the present invention. Using this formula, the change of weight ω under different factors can be explored. The results are as Figure 2 shown. It can be seen that the value of ω varies between 0.65 and 1.84.
[0035] Exploration of the relationship between 1.5ω and known information
[0036] Although the numerical change characteristics of ω have been understood through simulation data, it is difficult to find a specific value of ω in practice to improve the accuracy rate. This part will focus on exploring ω corresponding to different data. The basic goal is to predict ω using the known information of the data, so as to find the optimal replacement value. For pollutant data containing below the detection limit, <LOD%, the mean value E(X|X≥c) of data greater than the detection limit and the standard deviation sd(X|X≥o) of data greater than the detection limit are known information. Therefore, the relationship between these known information and ω is focused on exploring. The exploration results are shown in Figure 3 , from which it can be seen that ω decreases with the increase of the ratio sd(X|X≥c) / E(X|X≥c) and decreases with the increase of <LOD%. Most importantly, ω has an obvious relationship with these known information, which is an important discovery in the exploration of the present invention. But Figure 3 the presented relationship is a curve relationship. In reality, the quantitative description of the curve is diverse, which is not conducive to constructing a stable curve relationship. Therefore, an attempt is made to Figure 4 perform logarithmic transformation on the variables in Figure 4 in order to expect to construct a linear relationship. The results are as Figure 4Only the relationship when <LOD% is between 5% and 50% is presented. Apparently, after logarithmic transformation, the weight coefficient ω has a strong linear relationship with lnsa(X|X≥c) / E(X|X≥c), but there is a contribution from <LOD% in this relationship. This is another important discovery explored in the present invention. In order to find an expression for quantitatively describing this relationship, the applicant tried to introduce variables in a stepwise regression manner to compare the goodness of fit in different variable combinations, and finally selected the optimal equation, which is summarized as the following formula:
[0037] ω = (-229X1 - 8.29X2 - 1.44X1X2 + 1495) / 1000,
[0038] In the above formula, X1 = lnA,
[0039]
[0040] X2 is the proportional value of the test samples with pollutant concentration below LOD.
[0041] So far, through the onion - peeling exploration, using the above formula, environmental science workers can easily calculate the arithmetic mean with high precision when encountering data containing values below the detection limit.
[0042] Next, compare the advantages and disadvantages of LOD / 2 and ω*LOD / 2 in replacing the data below the detection limit to calculate the mean value. The present invention first generates log - normal random numbers with a sample size of 160. The specific parameter settings are described in Section 1.2. Then set <LOD% to vary from 5% to 50%. On this basis, obtain the LOD of the data, then replace the data with <LOD in the random numbers with LOD / 2, and then calculate the mean value; at the same time, according to the known information of the data, calculate ω according to the formula ω = (-229X1 - 8.29X2 - 1.44X1X2 + 1495) / 1000, replace the data with <LOD in the random numbers with ω*LOD / 2, and then calculate the mean value. Compare the advantages and disadvantages of the two methods according to the true mean value of the random numbers. The results are shown in Figure 5 , from Figure 5 it can be seen that the ω*LOD / 2 replacement method proposed by the present invention not only has high average accuracy but also is more stable. It can be said to be an excellent algorithm for calculating the mean value containing values below the detection limit.
[0043] Taking soil o-xylene pollutants as an example, the actual use of this invention will be elaborated. The applicant conducted an investigation on underground soil petroleum pollution at a certain abandoned petroleum storage site. The site has been in use for more than 20 years, and the main pollutant is petroleum hydrocarbons. The pollution depth is > 3m, and the deep soil pollution is very serious. The o-xylene lost through volatilization is negligible. O-xylene has an irritating effect on the human respiratory system and can damage the skin and mucous membranes. In addition, o-xylene can harm the human nervous system, and symptoms such as conjunctival and pharyngeal congestion, dizziness, nausea, vomiting, chest tightness, weakness in the limbs, confusion, and staggering gait may occur. In order to explore the actual concentration of o-xylene at this site, a total of 57 soil samples were collected using drilling equipment for analysis. The LOD is 0.0015 mg / kg. Among them, 41 samples are above the detection limit (LOD), and 16 samples are below the LOD. The samples below the detection limit do not have an obvious distribution pattern, which can be considered to be caused by the heterogeneity of the soil. Obviously, the average value of the pollutants cannot be directly calculated. Now, calculate the mean according to the steps provided by the present invention.
[0044] (1) First, calculate the proportion of pollutant data below the detection limit containing pollutants below the detection limit,
[0045] <LOD%=16 / 57*100=28.1%
[0046] (2) Second, calculate the mean (E(X|X≥c)) and standard deviation (sd(X|X≥c)) of the pollutant concentration greater than the detection limit of the pollutants,
[0047] The mean of the pollutant concentration greater than the detection limit = 8.11, and the standard deviation of the pollutant concentration greater than the detection limit = 19.78,
[0048] (3) ω=(-229ln(19.78 / 8.11)-8.29×28.1-1.44ln(19.78 / 8.11)×28.1+
[0049] 1495) / 1000=1.02,
[0050] (4) The replacement value ω*LOD / 2=1.02*0.0015 / 2=0.000765. Replace this value with the data below the detection limit. Finally, the obtained mean ([8.11×41+0.000765×16] / 57)=5.83mg / kg.
[0051] In addition, it should be understood that although this specification is described according to the embodiments, not each embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A method for estimating the average value of pollutant concentration, characterized in that, The specific steps are as follows: Obtain the detection data of the pollutant concentration in multiple detection samples, and replace the average value of the pollutant concentration below the LOD with ω*LOD / 2, where ω is obtained from the detection data of the pollutant concentration above the LOD; ω has a linear relationship with X1, and X1 = lnA, ω = (-229X1 - 8.29X2 - 1.44X1X2 + 1495) / 1000, where X2 is the proportion value of the detection samples with the pollutant concentration below the LOD.
2. The method for estimating the average value of pollutant concentration according to claim 1, characterized in that In the said detection samples, the proportion of the detection samples with the pollutant concentration below the LOD is 5% - 50%.
3. The method for estimating the average value of pollutant concentration according to claim 1, wherein, 0.65 ≤ ω ≤ 1.84.
Citation Information
Patent Citations
Bayesian framework-based food risk assessment method and device
CN114491414A