Ecological risk assessment method for organic pollutants in natural water body
Through machine learning and feature screening, quantitative structure-activity relationship toxicity prediction model is established, combined with risk entropy method and comprehensive scoring method, the problems of data scarcity and inaccurate evaluation in persistent organic pollutant assessment are solved, and the ecological risks of organic pollutants in water bodies are achieved efficiently.
Patent Information
- Application Number
- CN202510310777.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art has difficulty in effectively assessing the risks of persistent organic pollutants to aquatic ecosystems, especially due to the lack of sufficient aquatic biotoxicity data and the limitations of traditional assessment methods, resulting in inaccurate risk assessment.
A quantitative structure-activity relationship toxicity prediction model was established using machine learning and feature screening methods, combined with risk entropy method and comprehensive scoring method, toxicity data was collected and analyzed through databases, species sensitivity distribution model was established, predicted invalid concentrations were calculated, and the hazards of pollutants in water bodies were evaluated.
It improves the accuracy of evaluating the toxicity of persistent organic pollutants in seawater and freshwater, comprehensively evaluates the ecological risks of organic pollutants in water bodies, makes up for the shortcomings of traditional methods, and improves the accuracy and comprehensiveness of the assessment.
Smart Images

Figure CN120260718A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of ecological risk assessment of natural water bodies, and particularly to an ecological risk assessment method for organic pollutants in natural water bodies. Background Art
[0002] Persistent organic pollutants refer to a class of organic pollutants that are difficult to degrade in the natural environment, having persistence, mobility, and biological toxicity, and can accumulate along the food chain. Currently, common persistent organic pollutants include polychlorinated biphenyls (PCBs), polybrominated diphenyl ethers (PBDEs), and per- and polyfluoroalkyl substances (PFASs), etc.
[0003] In recent years, the number of newly added persistent organic pollutants globally has increased sharply, causing serious harm to the aquatic ecosystem. There is an urgent need to carry out the work of water ecological risk assessment of pollutants. However, based on the species sensitivity distribution (SSD) method, at least 3 phyla and 8 families of aquatic biological toxicity data are required to evaluate the predicted no-effect concentration (PNEC) value of a compound. Currently, the toxicity data of persistent organic pollutants to aquatic organisms are very scarce. Conducting systematic and comprehensive toxicity tests requires a long period and high cost. At the same time, it is very difficult to ensure the survival rate of organisms under laboratory conditions. The risk assessment requirements of pollutants cannot be met only through experimental methods.
[0004] In view of this, it is necessary to provide a new technical solution to solve the above problems. Summary of the Invention
[0005] To solve the above technical problems, the present application provides an ecological risk assessment method for pollutants in water bodies, which can effectively supplement the toxicity data of pollutants, effectively evaluate the toxicity of persistent organic pollutants in seawater and freshwater, and then effectively evaluate the ecological risk of organic pollutants in water bodies.
[0006] An ecological risk assessment method for organic pollutants in natural water bodies includes:
[0007] S1. Establish a quantitative structure-activity relationship toxicity prediction model of species based on machine learning and feature screening methods;
[0008] S2. Use the quantitative structure-activity relationship model to predict and obtain the biological toxicity data in the water body to be predicted, and establish a species sensitivity distribution model of pollutants in the water body based on this;
[0009] S3. Calculate the predicted no-effect concentration using the species sensitivity distribution model;
[0010] S4. Based on database retrieval, acute toxicity experiments, and interspecies relationship prediction models, collect, predict, and supplement the toxicity data of species in water bodies, and evaluate the harmfulness of pollutants in water bodies using the risk entropy method and the comprehensive scoring method respectively.
[0011] Preferably, step S1 includes:
[0012] S11. Collect toxicity data using existing database and literature database data, and preprocess the collected data;
[0013] S12. Randomly divide the preprocessed data into a training set and a validation set;
[0014] S13. Establish several machine learning algorithm models for several species as quantitative structure-activity relationship toxicity prediction models, and train the established models;
[0015] S14. Verify and evaluate the goodness of fit, robustness, and prediction ability of each trained quantitative structure-activity relationship toxicity prediction model, and select the optimal model as the final quantitative structure-activity relationship toxicity prediction model.
[0016] Preferably, when training the established quantitative structure-activity relationship toxicity prediction model, it also includes feature screening for each established quantitative structure-activity relationship toxicity prediction model, and using the screened variables for model training, where the feature screening includes:
[0017] S131. Calculate the information gain of each variable and sort them;
[0018] S132. Eliminate the variable ranked last according to the feature importance ranking to obtain a new feature set;
[0019] S132. Re-execute the processes of step 131 and step 132 with the new feature set until the last 2 variables remain as the screened variables.
[0020] Preferably, when collecting toxicity data using existing database and literature database data, use 1 - 4 days as the exposure time, use acute toxicity data EC 50 and LC 50 as the toxicity endpoints of the toxicity data, and use the growth rate, activity inhibition rate, or lethality rate as the toxicity effect.
[0021] Preferably, when collecting toxicity data using existing database and literature data, if there are multiple toxicity values for an organic pollutant in the same species, select the toxicity value of the species at the sensitive period; if there are multiple toxicity values within no more than one order of magnitude for the same species at the sensitive period, calculate the geometric mean of the multiple toxicity values; if there are multiple toxicity values for the same species at the sensitive period that exceed one order of magnitude, analyze whether the experimental processes for obtaining each toxicity value conform to the standard toxicity test method. On the premise that all experimental processes conform to the standard toxicity test method, select the data with the most standard experimental process as the species toxicity value for collection; if there are multiple toxicity values for the same species at the sensitive period that exceed one order of magnitude, analyze whether the experimental processes for obtaining each toxicity value conform to the standard toxicity test method. If all experimental processes do not conform to the standard toxicity test method, delete all such data.
[0022] Preferably, the calculation formula for the geometric mean of multiple toxicity values is:
[0023]
[0024] In the formula, SMAV refers to the toxicity value with the same effect, ATV refers to the acute toxicity value, i refers to a certain species, k is the type of acute toxicity effect, and m is the number of SMAV.
[0025] Preferably, step S3 includes:
[0026] S31. Arrange the selected species in ascending order of toxicity and calculate the cumulative probability of the species:
[0027]
[0028] In the formula, P is the cumulative probability, i is the ranking of the species toxicity, and n is the total number of toxicity data;
[0029] S31. Using the cumulative probability as the ordinate and the logarithm of the EC 50 or LC 50 of each species as the abscissa, fit using different models. Finally, perform a KS test on the fitted model, calculate the root mean square error and the coefficient of determination, and select the optimal fitted model;
[0030] S31. Obtain HC5 in the optimal fitted curve, extrapolate PNEC based on HC5, and obtain the prediction invalid concentration fixed value formula:
[0031]
[0032] Among them, PNEC is the predicted no-effect concentration; HC5 is the concentration corresponding to the cumulative probability of 5% in the species sensitivity distribution curve; AF is the assessment factor, which is comprehensively determined according to the amount of data used to derive PNEC, the coverage of test organisms, and the fitting distribution of data, etc. Its value range is 2 - 5; when the number of species is greater than 20, the value of the assessment factor is 2; when the number of species is no more than 20, the value of the assessment factor is 3.
[0033] Preferably, when evaluating the harmfulness of pollutants in water bodies using the comprehensive scoring method, it includes:
[0034] S41. Select four indicators: the persistence, bioaccumulation, toxicity, and concentration of the pollutant;
[0035] S42. Assign a weight to each indicator according to the importance of each indicator in ecological risk assessment;
[0036] S43. After determining the weights, standardize each indicator to eliminate the influence caused by differences in indicator dimensions and value ranges;
[0037] S44. Multiply each standardized indicator by its corresponding weight, and then sum to obtain the comprehensive score.
[0038] Preferably, when standardizing each indicator, it includes:
[0039] Perform logarithmic transformation on all data;
[0040] Use the min - max regularization method to standardize each indicator.
[0041] Preferably, when evaluating the harmfulness of pollutants in water bodies using the risk entropy method, a risk entropy less than 0.01 indicates no risk; a risk entropy greater than 0.01 and less than 0.1 indicates low risk; a risk entropy greater than 0.1 and less than 1 indicates medium risk; a risk entropy greater than 1 indicates high risk;
[0042] The expression of risk entropy is:
[0043]
[0044] In the formula, RQ s is the risk entropy; MEC is the maximum concentration of the measured or predicted pollutant in the environment; PNEC is the predicted no - effect concentration.
[0045] Compared with the prior art, the present application has at least the following beneficial effects:
[0046] 1. The present invention uses machine learning methods for ecological risk assessment, which can effectively evaluate the toxicity of persistent organic pollutants in seawater and freshwater, and thus can effectively evaluate the ecological risk of organic pollutants in water bodies.
[0047] 2. The present invention predicts and obtains the biotoxicity data in the water body to be predicted by using a quantitative structure-activity relationship model, providing sufficient biotoxicity data for ecological risk, and improving the accuracy of the ecological risk assessment of persistent organic pollutants in seawater and freshwater.
[0048] 3. The present invention respectively uses the risk entropy method and the comprehensive scoring method to evaluate the harmfulness of pollutants in the water body, and uses the comprehensive scoring method to make up for the deficiencies of the risk entropy method, further improving the accuracy of the ecological risk assessment of persistent organic pollutants in seawater and freshwater. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Some specific embodiments of the present invention will be described in detail hereinafter with reference to the accompanying drawings in an exemplary but not restrictive manner. The same reference numerals in the drawings denote the same or similar components or parts. Those skilled in the art should understand that these drawings are not necessarily drawn to scale.
[0050] In the drawings:
[0051] Figure 1 is a schematic diagram of the overall process of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0053] In the risk assessment of organic pollutants, especially when calculating the predicted no-effect concentration (PNEC) of pollutants, there is a lack of sufficient local species-related data, which may lead to overprotection or underprotection problems in the risk assessment of this natural water body area.
[0054] Solving the above problems requires a large amount of risk assessment-related data. However, for hundreds of thousands of chemicals on the market, only a very small number (<0.5%) have relevant toxicity data, severely limiting the assessment and prediction of their environmental exposure, harmfulness, and risks. Therefore, constructing a quantitative structure-activity relationship toxicity prediction model can efficiently obtain the toxicity parameters of chemical substances and provide a scientific basis for the source control of organic pollutants.
[0055] In addition, only using the traditional risk quotient (RQs) assessment method may ignore the persistence and bioaccumulation of new pollutants, thus failing to comprehensively evaluate the ecological risks of new pollutants. By using the risk quotient method and the comprehensive scoring method to evaluate the harmfulness of pollutants in water respectively, the deficiency of the risk quotient method can be made up by the comprehensive scoring method, further improving the accuracy of the ecological risk assessment of persistent organic pollutants in seawater and freshwater.
[0056] As Figure 1 shown, an ecological risk assessment method for organic pollutants in natural water bodies includes the following steps:
[0057] S1. Establish a quantitative structure-activity relationship (QSAR) toxicity prediction model for species based on machine learning and feature screening methods.
[0058] Specifically, it includes:
[0059] S11. Collect toxicity data using existing database and literature data, and preprocess the collected data.
[0060] Specifically, when collecting toxicity data using existing database and literature data, use 1 - 4 days as the exposure time, use acute toxicity data EC 50 and LC 50 as the toxicity endpoints of toxicity data, and use the growth rate, activity inhibition rate or lethality rate as the toxicity effect.
[0061] It should be noted that LC 50 is the median lethal concentration, and EC 50 is the median effect concentration.
[0062] In addition, when collecting toxicity data using existing database and literature data, if there are multiple toxicity values of the selected organic pollutant under the same species, select the species toxicity value in the sensitive period; if there are multiple toxicity values of the same species in the sensitive period that do not exceed one order of magnitude, calculate the geometric mean of the multiple toxicity values; if there are multiple toxicity values of the same species in the sensitive period that exceed one order of magnitude, analyze whether the experimental processes for obtaining each toxicity value meet the standard toxicity test method. On the premise that all experimental processes meet the standard toxicity test method, select the data with the most standard experimental process as the species toxicity value for collection; if there are multiple toxicity values of the same species in the sensitive period that exceed one order of magnitude, analyze whether the experimental processes for obtaining each toxicity value meet the standard toxicity test method. If all experimental processes do not meet the standard toxicity test method, delete all such data.
[0063] In this embodiment, the toxicity data required for constructing the QSAR toxicity prediction of aquatic organisms is collected from the ECOTOX database, CNKI and WOS database.
[0064] The calculation formula for the geometric mean of multiple toxicity values is as follows:
[0065]
[0066] In the formula, SMAV refers to the toxicity value with the same effect, ATV refers to the acute toxicity value, i refers to a certain species, k is the type of acute toxicity effect, and m is the number of SMAV.
[0067] S12. Randomly divide the preprocessed data into a training set and a validation set.
[0068] S13. Establish several machine learning algorithm models for several species as quantitative structure-activity relationship toxicity prediction models, and train the established models.
[0069] When training the established quantitative structure-activity relationship toxicity prediction models, it also includes feature screening for each established quantitative structure-activity relationship toxicity prediction model, and using the screened variables for model training. Among them, the feature screening includes:
[0070] S131. Calculate the information gain of each variable and sort them.
[0071] S132. Eliminate the variable ranked last according to the feature importance ranking to obtain a new feature set.
[0072] S132. Re-execute the processes of step 131 and step 132 with the new feature set until the last 2 variables are left as the screened variables.
[0073] S14. Verify and evaluate the goodness of fit, robustness, and prediction ability of each trained quantitative structure-activity relationship toxicity prediction model, and select the optimal model as the final quantitative structure-activity relationship toxicity prediction model.
[0074] In this embodiment, based on 12 datasets of the Tox21 Challenge, 76,769 models were constructed using 7 machine learning algorithms including decision tree (DT), naive Bayes (NB), k-nearest neighbor algorithm (KNN), support vector machine (SVM), multi-layer perceptron (MLP), random forest (RF), and extreme gradient boosting (XGB), as well as the feature screening method. The area under the receiver operating characteristic curve (AUC) of the specified validation set in the Tox21 Challenge was selected as the index for evaluating the final ranking. The results show that in the 12 datasets and 7 machine learning algorithms, the AUC of the models after feature screening is greater than or equal to the AUC of the models before feature screening, indicating that the feature screening method is beneficial to improving the prediction performance of the models.
[0075] In addition, the performance of the optimal models of different algorithms was compared in 12 datasets, and it was found that the RF algorithm performed best in the AhR, AR, ER, PPAR-g, ARE, ATAD5, HSE, p53, and NR-LBD datasets, and the XGB algorithm performed best in the Aromatase, ER, ER-LBD, and MMP datasets. The predictive abilities of other algorithms were lower than those of these two algorithms. Therefore, the RF algorithm model and the XGB algorithm model were used as the final quantitative structure-activity relationship toxicity prediction models.
[0076] Model performance evaluation showed that the XGB algorithm and the feature screening method were more compatible, which could significantly improve the predictive ability of the model. The coefficient of determination increased by an average of 22%, which was 2 percentage points higher than that of the simple RF algorithm model (20%) when performing toxicity assessment. In terms of computational complexity, since the feature screening method requires n times of training (n is the number of variables in the model), the computational complexity of the model also increases by n times. When the amount of data is too large, it will greatly extend the training time of the model. The RF algorithm model performs more stably before and after feature screening, so it is more suitable for situations with insufficient computing resources. Generally speaking, using the RF algorithm, the XGB algorithm, and the feature screening method can improve the prediction accuracy of the QASR model, thereby reducing the difference between the prediction result and the actual value and improving the accuracy of ecological risk assessment.
[0077] In addition, in this implementation, the Euclidean distance (D E ) was used to evaluate the application domain of the QSAR model, and the formula was:
[0078]
[0079] where x is the row vector of the data point, μ is the mean of x, and the superscript T is the transpose. If the distance between the pollutant and the center point of the training set exceeds the maximum distance of the training set, then the pollutant is considered to be outside the application domain of the model.
[0080] S2. Use the quantitative structure-activity relationship model to predict and obtain the biotoxicity data in the water body to be predicted, and establish a species sensitivity distribution (SSD) model of pollutants in the water body based on this.
[0081] The species sensitivity distribution method is a statistical extrapolation method based on extrapolating the toxicological endpoints of a single species to the biological community or ecosystem level. Its principle is that in a complex ecosystem, the sensitivities of different species to a certain environmental stress follow a specific probability distribution, including the log-logistic, log-normal, Gumbel, Burr, and Weibull distributions. The corresponding cumulative distribution function (CDF) is as follows:
[0082]
[0083] Wherein, x is the value of the toxicity data of each species; α is the scale parameter; β is the shape parameter.
[0084] S3. Calculate the predicted no-effect concentration using the species sensitivity distribution model.
[0085] Specifically, it includes:
[0086] S31. Arrange the selected species in ascending order of toxicity, and calculate the cumulative probability of the species:
[0087]
[0088] Wherein, P is the cumulative probability, i is the toxicity ranking of the species, and n is the total number of toxicity data;
[0089] S32. Taking the cumulative probability as the ordinate and the logarithm of the EC 50 or LC 50 of each species as the abscissa, use different models for fitting, finally perform the KS test on the fitting model, calculate the root mean square error and the coefficient of determination, and select the optimal fitting model;
[0090] S33. Obtain the HC5 in the optimal fitting curve, and extrapolate the PNEC according to the HC5 to obtain the formula for the predicted no-effect concentration value:
[0091]
[0092] Among them, PNEC is the predicted no-effect concentration; HC5 is the concentration corresponding to the cumulative probability of 5% in the species sensitivity distribution curve; AF is the assessment factor, which is determined comprehensively according to the amount of data used to derive the PNEC, the coverage range of the test organisms, and the fitting distribution of the data, etc. Its value range is 2 - 5; when the number of species is greater than 20, the value of the assessment factor is 2; when the number of species is not greater than 20, the value of the assessment factor is 3.
[0093] S4. Based on database retrieval, acute toxicity experiments and interspecies relationship prediction models, collect, predict and supplement the species toxicity data in water bodies, and evaluate the harmfulness of pollutants in water bodies using the risk entropy method and the comprehensive scoring method respectively.
[0094] Specifically, when evaluating the harmfulness of pollutants in water bodies using the risk entropy method, a risk entropy less than 0.01 indicates no risk; a risk entropy greater than 0.01 and less than 0.1 indicates low risk; a risk entropy greater than 0.1 and less than 1 indicates medium risk; a risk entropy greater than 1 indicates high risk;
[0095] The expression of the risk entropy is:
[0096]
[0097] In the formula, RQ s is the risk entropy; MEC is the maximum concentration of the measured or predicted pollutant in the environment; PNEC is the predicted no-effect concentration.
[0098] Furthermore, when using the comprehensive scoring method to evaluate the harmfulness of pollutants in water bodies, it includes:
[0099] S41. Select four indicators: the persistence, bioaccumulation, toxicity, and concentration of the pollutant.
[0100] S42. Assign a weight to each indicator according to its importance in the ecological risk assessment.
[0101] S43. After determining the weights, standardize each indicator to eliminate the influence caused by differences in indicator dimensions and numerical ranges.
[0102] Among them, when standardizing each indicator, first perform logarithmic transformation on all data, and then use the min-max regularization method to standardize each indicator:
[0103]
[0104] Among them, X nom refers to the normalized data, X refers to the data before normalization, X min refers to the minimum value of the data, X max refers to the maximum value of the data.
[0105] S44. Multiply each standardized indicator by its corresponding weight, and then sum to obtain the comprehensive score.
[0106] The advantage of the method based on RQs is that there are corresponding reference standards, which can quickly evaluate the current level of ecological risk. However, it only considers the toxicity and concentration parameters and lacks comprehensive consideration of other factors. In contrast, the comprehensive scoring method has the advantage of considering multiple ecological risk factors, exploring the correlation and complementarity between various factors, and improving the reliability and accuracy of the evaluation results.
[0107] To comprehensively consider subjective and objective evaluation methods, the AHP-Critic comprehensive scoring method is adopted in this method to calculate the weights of indicators. Among them, the expert scores of the AHP method are derived from a questionnaire survey conducted by experts (n = 11) in environmental fields such as hydrogeology, chemistry, toxicity, management, investigation, pollution prevention, and groundwater control in existing literature. The Critic method reflects the weight coefficients between them by calculating the standard deviation between indicators. Finally, the weight results of the two parts are averaged as the final weight, and the calculation process is implemented through SPSS-PRO. The results of the expert questionnaire survey are shown in Table 1.
[0108] Table 1 Results of the expert questionnaire survey
[0109] expert Concentration (%) Persistence (%) Bioaccumulation (%) Toxicity (%) 1 13.82% 19.02% 29.43% 37.72% 2 41.11% 4.08% 3.95% 50.86% 3 58.29% 5.65% 5.65% 30.42% 4 39.85% 9.01% 2.28% 48.86% 5 63.24% 16.21% 7.43% 13.12% 6 26.61% 36.81% 20.84% 15.74% 7 38.38% 9.20% 9.20% 43.22% 8 12.53% 63.04% 9.86% 14.58% 9 11.86% 3.85% 1.69% 82.59% 10 60.82% 23.38% 8.53% 7.27% 11 5.77% 33.11% 12.90% 48.22% Average 33.84% 20.31% 10.16% 35.69%
[0110] In the present invention, the RQs method and the comprehensive scoring method are combined to evaluate organic pollutants in natural waters, which can make up for the deficiencies of the RQs method, and then screen out some compounds with potential risks, improving the accuracy of ecological risk assessment.
[0111] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0112] It should be noted that the terms "first", "second", etc. in the description and claims of the present application and the above drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein.
[0113] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An ecological risk assessment method for organic pollutants in natural water bodies, characterized in that, Including: S1. Establish a quantitative structure-activity relationship toxicity prediction model for species based on machine learning and feature screening methods; S2. Use the quantitative structure-activity relationship model to predict and obtain the biotoxicity data in the water body to be predicted, and establish a species sensitivity distribution model for pollutants in the water body based on this; S3. Calculate the predicted no-effect concentration using the species sensitivity distribution model; S4. Based on database retrieval, acute toxicity experiments, and interspecies relationship prediction models, collect, predict, and supplement the species toxicity data in the water body, and evaluate the harmfulness of pollutants in the water body using the risk entropy method and comprehensive scoring method respectively.
2. The ecological risk assessment method according to claim 1, wherein, Step S1 includes: S11. Collect toxicity data using existing database and literature database data, and preprocess the collected data; S12. Randomly divide the preprocessed data into a training set and a validation set; S13. Establish several machine algorithm models for several species as quantitative structure-activity relationship toxicity prediction models, and train the established models; S14. Verify and evaluate the goodness of fit, robustness, and prediction ability of each trained quantitative structure-activity relationship toxicity prediction model, and select the optimal model as the final quantitative structure-activity relationship toxicity prediction model.
3. The ecological risk assessment method according to claim 2, wherein When training the established quantitative structure-activity relationship toxicity prediction model, it also includes feature screening for each established quantitative structure-activity relationship toxicity prediction model, and using the screened variables for model training. Among them, feature screening includes: S131. Calculate the information gain of each variable and sort them; S132. Eliminate the variable ranked last according to the feature importance ranking to obtain a new feature set; S132. Re-execute the processes of step 131 and step 132 with the new feature set until the last 2 variables remain as the screened variables.
4. The ecological risk assessment method according to claim 3, wherein When collecting toxicity data using existing database and literature data, 1-4 days is used as the exposure time, and acute toxicity data EC 50 and LC 50 are used as the toxicity endpoints of the toxicity data, and the growth rate, activity inhibition rate or lethality rate is used as the toxicity effect.
5. The ecological risk assessment method according to claim 4, wherein When collecting toxicity data using existing database and literature database data, if there are multiple toxicity values for the selected organic pollutant under the same species, select the species toxicity value in the sensitive period; if there are multiple toxicity values within no more than one order of magnitude for the same species in the sensitive period, calculate the geometric mean of the multiple toxicity values; if there are multiple toxicity values for the same species in the sensitive period that exceed one order of magnitude, analyze whether the experimental processes for obtaining each toxicity value conform to the standard toxicity test method. On the premise that the experimental processes all conform to the standard toxicity test method, select the data with the most standard experimental process as the species toxicity value for collection; if there are multiple toxicity values for the same species in the sensitive period that exceed one order of magnitude, analyze whether the experimental processes for obtaining each toxicity value conform to the standard toxicity test method. If the experimental processes do not conform to the standard toxicity test method, delete all such data.
6. The ecological risk assessment method according to claim 5, wherein The calculation formula for the geometric mean of multiple toxicity values is: In the formula, SMAV refers to the same-effect toxicity value, ATV refers to the acute toxicity value, i refers to a certain species, k is the type of acute toxicity effect, and m is the number of SMAV.
7. The ecological risk assessment method according to claim 1, wherein Step S3 includes: S31. Sort the selected species in ascending order of toxicity and calculate the cumulative probability of the species: In the formula, P is the cumulative probability, i is the species toxicity ranking, and n is the total number of toxicity data; S31. Using the cumulative probability as the ordinate and the logarithm of the EC 50 or LC 50 of each species as the abscissa, fitting is performed using different models. Finally, the KS test is carried out on the fitting model, the root mean square error and the coefficient of determination are calculated, and the optimal fitting model is selected; S31. Obtain HC5 in the optimal fitting curve, and extrapolate PNEC based on HC5 to obtain the prediction formula for the threshold value of the no-effect concentration: Among them, PNEC is the predicted no-effect concentration; HC5 is the concentration corresponding to the cumulative probability of 5% in the species sensitivity distribution curve; AF is the assessment factor, which is comprehensively determined according to the amount of data used for deriving PNEC, the coverage of test organisms, and the fitting distribution of data, etc. Its value range is 2 - 5; when the number of species is greater than 20, the value of the assessment factor is 2; when the number of species is not greater than 20, the value of the assessment factor is 3.
8. The ecological risk assessment method according to claim 1, wherein When evaluating the harmfulness of pollutants in water bodies using the comprehensive scoring method, it includes: S41. Select four indicators: the persistence, bioaccumulation, toxicity, and concentration of the pollutant; S42. Assign a weight to each indicator according to its importance in the ecological risk assessment; S43. After determining the weights, standardize each indicator to eliminate the influence caused by differences in indicator dimensions and numerical ranges; S44. Multiply each standardized indicator by its corresponding weight, and then sum them up to obtain the comprehensive score.
9. The ecological risk assessment method according to claim 8, characterized in that, When standardizing each indicator, it includes: Perform logarithmic transformation on all data; Use the min-max regularization method to standardize each indicator.
10. The ecological risk assessment method according to claim 8 or 9, characterized in that, When evaluating the harmfulness of pollutants in water bodies using the risk entropy method, a risk entropy less than 0.01 indicates no risk; a risk entropy greater than 0.01 and less than 0.1 indicates low risk; A risk entropy greater than 0.1 and less than 1 indicates medium risk; a risk entropy greater than 1 indicates high risk; The expression of the risk entropy is: where RQ s is the risk entropy; MEC is the maximum concentration of the measured or predicted pollutant in the environment; PNEC is the predicted no-effect concentration.
Citation Information
Cited By
Bioaccumulation pollutant chronic toxicity prediction system and method
CN120822659A
Water pollution event index screening method and device, electronic equipment and storage medium
CN121032340A