Surface water organophosphate pollution ecological risk prediction method and system
By combining random forest and Monte Carlo method for eliminating uninformed variables with cross-validation, a quantitative model for surface water organophosphate concentration was constructed, which solved the problem of accurately predicting surface water organophosphate concentration and ecological risk, and achieved more efficient pollution prevention and control.
Patent Information
- Application Number
- CN202510962838.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-07-14
AI Technical Summary
Existing technologies are insufficient to accurately predict the concentration and ecological risks of organophosphates in surface water in different regions, resulting in inadequate accuracy and efficiency in pollution prevention and control.
A Monte Carlo Uninformation Variable Elimination (MC-UVE-RF) method based on random forest was used to screen socioeconomic and environmental variables. By combining cross-validation and parameter search, a random forest concentration quantitative model was constructed. The model parameters were optimized through Monte Carlo simulation to achieve accurate prediction of organophosphate concentration and quantify the ecological risk level.
It improves the accuracy of predicting organophosphate concentrations in surface water and the precision of ecological risk assessment, enabling better identification of high-risk areas and enhancing the accuracy and efficiency of pollution prevention and control.
Smart Images

Figure CN120954552A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of environmental monitoring technology, and in particular relates to a method and system for predicting the ecological risk of surface water organic phosphate pollution. Background Technology
[0002] Organophosphates, as a novel pollutant, are widely used in industrial and consumer products as flame retardants, plasticizers, and additives. With the phasing out of brominated flame retardants, the production and application of organophosphates have surged. Studies have shown that organophosphates are associated with numerous adverse human health outcomes, including cancer, digestive disorders, neurological problems, and reproductive disorders. Organophosphates are primarily incorporated into materials through physical mixing rather than chemical bonding, leading to their continuous release into the aquatic environment during production, use, and disposal via leaching, volatilization, and abrasion. Therefore, organophosphates are widely distributed in rivers, lakes, and even source waters, but their concentration and composition vary significantly across different regions and water bodies. This invention employs an efficient method to predict the concentration and risk of organophosphates, enabling targeted control in high-risk areas and improving the accuracy and efficiency of pollution prevention and control. Summary of the Invention
[0003] In view of this, the present invention aims to overcome the shortcomings of the above-mentioned problems in the prior art and proposes a method and system for predicting the ecological risk of surface water organophosphate pollution.
[0004] To achieve the above objectives, the technical solution of the present invention is implemented as follows:
[0005] The first aspect of this invention proposes a method for predicting the ecological risk of surface water organophosphate pollution, comprising the following steps:
[0006] S1: Obtain the dataset of observed organic phosphate concentrations in the target area, as well as socioeconomic and environmental variable data. Preprocess the data and randomly divide it into training and test sets.
[0007] S2: Socioeconomic and environmental variables were screened using the Monte Carlo Uninformation Variable Elimination Method (MC-UVE-RF) based on random forest to remove uninformation variables and select variables that have a significant impact on concentration, which were then used as the dataset for model construction.
[0008] S3: Construct a random forest concentration quantitative model using the model building dataset and the observed sample organic phosphate concentration dataset. Employ a combination of cross-validation and parameter search. Optimize the key parameters of the random forest model based on model evaluation metrics to construct the optimal model. Determine the accuracy of the prediction using model evaluation metrics on the model test set. If the prediction is deemed inaccurate, rerun MC-UVE-RF and adjust the dummy variable ratio or PIP threshold until the prediction is accurate. Determine the model parameters and obtain a converged organic phosphate concentration quantitative model.
[0009] S4: Input the target region model construction variable dataset into the converged organophosphate concentration quantitative model to predict the organophosphate concentration in the target region;
[0010] S5: Use predicted values of organophosphate concentrations to quantify ecological risks, identify ecological risk levels, and project these levels onto target areas.
[0011] As a further option, in S1, the socio-economic and environmental variable data include conventional hydrological data, geographical data, climate data, soil data, social and economic data, etc.
[0012] As a further step, in S1, the preprocessing includes: handling missing values in the observed sample organophosphate concentration data and socioeconomic and environmental data, cleaning outliers, and performing logarithmic transformation on the data.
[0013] As a further measure, in S1, the data is randomly divided into a training set and a test set as follows: 70% of all sample data is randomly selected as the training set, and the remaining 30% is the test set. To ensure that the concentration of the training set covers all predicted samples, the three samples with the highest and lowest concentrations are manually added to the training set.
[0014] As a further solution, S2 specifically includes the following steps:
[0015] S21: Generate virtual random uninformative variables with 5 times the number of variables, mix them with real variables, and construct a random forest model by randomly sampling a certain number of variables using Monte Carlo methods.
[0016] S22: Calculate the random forest importance index of variables using the random forest model, and calculate the posterior inclusion probability (PIP) of each variable through Monte Carlo simulation based on this, and set the cutoff threshold.
[0017] S23: Based on the PIP and cutoff threshold, eliminate variables with no information and screen out socioeconomic and environmental variables that have a significant impact on concentration;
[0018] S24: Output the retained socioeconomic and environmental variables as the model building dataset.
[0019] As a further solution, in step S22, the importance of the indicator variable in the random forest is calculated as follows:
[0020]
[0021] Among them, MSE original This is expressed as the root mean square error of the original model. The mean squared error calculated after shuffling the j-th variable is %lncMSE, where %lncMSE represents the percentage increase in MSE after shuffling the eigenvalue, n represents the sample size, and y... i Let y represent the actual concentration of the i-th organophosphate ester. i ' represents the predicted concentration of the i-th organophosphate ester, y i (j) The predicted concentration of organophosphates is obtained by shuffling the j-th variable.
[0022] As a further solution, in S22, the posterior probability and cutoff threshold are calculated using the following formula:
[0023]
[0024] cutoff = k × max[abs(PIP)] noise )]
[0025] PIP j Let M represent the posterior inclusion probability of the j-th variable, and M represent the number of Monte Carlo simulations. This represents the random forest importance index for the j-th variable in the m-th sample. I(·) is the indicator function; if the condition within the parentheses is met, I = 1; otherwise, I = 0. It is used to count the effective occurrences. `cutoff` is the cutoff threshold, `k` is the adjustment coefficient, and `PIP` is the key factor. noise This represents the posterior inclusion probability score of an uninformative variable, and abs(·) is the absolute value function.
[0026] As a further approach, in S23, the method for screening socioeconomic and environmental variables that have a significant impact on concentration is to retain the posterior inclusion probability (PIP). j Socioeconomic and environmental variables exceeding the cutoff threshold were used as model building variables, and posterior inclusion probabilities (PIPs) were removed. j Uninformative variables below the cutoff threshold.
[0027] As a further option, the evaluation index mentioned in S3 is calculated as follows:
[0028]
[0029] Among them, RMSE is expressed as the root mean square error, and R 2 is expressed as the coefficient of determination, n is expressed as the number of samples, and y i is expressed as the actual value of the i-th organophosphate concentration, and y i ′ is expressed as the predicted value of the i-th organophosphate concentration, and is expressed as the average value of the actual values of the organophosphate concentrations.
[0030] As a further solution, the key parameters of the model described in S3 include: the number of decision trees, the maximum depth of a single tree, etc.
[0031] As a further solution, the method for determining the model parameters described in S3 is: constructing a parameter search space, selecting stratified k-fold cross-validation in combination with the task type to ensure balanced data distribution; then traversing the parameter combinations through grid search or random search, the former is suitable for precise search in a small range, and the latter is suitable for rough search on a large scale, and the model performance is evaluated based on the model evaluation index; finally, the optimal parameter combination is selected based on the cross-validation results, and the optimal model is constructed and its generalization ability is verified on the test set.
[0032] As a further solution, the criterion for judging accurate prediction described in S3 is: the relative root mean square error of the optimal model test set is less than the preset threshold and the coefficient of determination is greater than the preset threshold.
[0033] As a further solution, in S5, the ecological risk calculation method is to evaluate the ecological risk through the risk quotient (RQ) and divide it into levels:
[0034]
[0035] Among them, NOEC is the observed no-effect concentration, LOEC is the lowest observed effect concentration, ChV is the chronic no-effect value, which is estimated using ECOSAR software to ensure the consistency of the environmental risk assessment standards, RAF is the risk assessment factor, MEC is the predicted environmental concentration, and PNEC is the predicted no-effect concentration;
[0036] The estimated risk levels are divided into four levels: no obvious risk (RQ < 0.01), low risk (RQ < 0.1), medium risk (0.1 < RQ < 1), and high risk (RQ > 1).
[0037] The second aspect of the present invention provides a surface water organophosphate pollution ecological risk prediction system, including
[0038] a dataset acquisition module: obtaining the dataset of the observed sample organophosphate concentrations and environmental variable data in the target area, preprocessing the data, and randomly dividing the data into a training set and a test set;
[0039] Variable selection module: Socioeconomic and environmental variables are selected using the Monte Carlo Uninformation Variable Elimination Method (MC-UVE-RF) based on random forest to eliminate uninformation variables and select variables that have a significant impact on concentration, which are then used as the model construction dataset;
[0040] Model building module: A random forest concentration quantification model is built using the model building dataset and the observed sample organic phosphate concentration dataset. A combination of cross-validation and parameter search is used to optimize the key parameters of the random forest model according to the model evaluation index, build the optimal model, and judge the accuracy of the prediction by the model evaluation index of the model test set. If the prediction is determined to be inaccurate, MC-UVE-RF is rerun and the proportion of dummy variables or PIP threshold is adjusted until the prediction is accurate. The model parameters are determined and a converged organic phosphate concentration quantification model is obtained.
[0041] Concentration prediction module: Inputs the target region model construction variable dataset into the converged organophosphate concentration quantitative model to predict the organophosphate concentration in the target region;
[0042] Environmental and ecological risk identification module: Quantifies ecological risk by using predicted values of organophosphate concentrations, identifies ecological risk levels, and projects the ecological risk levels onto target areas.
[0043] Compared with existing technologies, the ecological risk prediction method and system for surface water organophosphate pollution described in this invention has the following advantages:
[0044] This invention incorporates Monte Carlo method-no-information variable elimination-random forest importance analysis into the variable selection stage and combines it with random forest to predict the ecological risk of organophosphate pollution in surface water. The combination of MC-UVE and random forest importance analysis distinguishes real signals from noise by simulating dummy variables, making it more suitable for complex environmental data than traditional correlation analysis. It better addresses the challenges of high-dimensional noise, small sample sizes, and nonlinear relationships in environmental data. Combined with the nonlinear modeling capabilities of the random forest model, a complete chain of "highly robust feature selection + strong prediction model" is formed, improving the accuracy of predicting organophosphate concentrations in surface water. Attached Figure Description
[0045] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings:
[0046] Figure 1 This is a schematic diagram of the process for predicting the ecological risk of surface water organophosphate pollution in Embodiment 1 of the present invention.
[0047] Figure 2This is an importance graph of random forest variables in Embodiment 1 of the present invention;
[0048] Figure 3 This is a ten-fold cross-validation diagram from Embodiment 1 of the present invention. Detailed Implementation
[0049] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other.
[0050] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined with "first," "second," etc., may explicitly or implicitly include one or more of that feature. In the description of this invention, unless otherwise stated, "a plurality of" means two or more.
[0051] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art will understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0052] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0053] Example 1:
[0054] See Figure 1 The purpose of this embodiment is to provide a method for predicting the ecological risk of organophosphate pollution in surface water, aiming to predict the concentration and ecological risk of organophosphate pollution in surface water. The specific prediction process is as follows: Figure 1 As shown:
[0055] Step S1: Obtain the dataset of observed organic phosphate concentrations and environmental variable data in the target area, and preprocess the data by randomly dividing it into training and test sets.
[0056] First, it is necessary to obtain the concentration data of the target organophosphates at the observation points. In this example, TCEP and EHDPP are selected as the target organophosphates based on the biotoxicity and application popularity of various types of organophosphates.
[0057] This example acquires environmental data corresponding to the observation points. It utilizes environmental data provided by the HydroATLAS database, including but not limited to conventional hydrological data, geographic data, climate data, soil data, and human data. Hydrological data includes, but is not limited to, water retention time, natural flow, lake volume, river area, and groundwater depth. Meteorological data includes, but is not limited to, temperature and precipitation. Geographic data includes, but is not limited to, altitude, terrain slope, and river gradient. Soil data includes, but is not limited to, soil clay content, soil organic carbon content, soil erosion, total soil phosphorus content, and total soil nitrogen content.
[0058] The data on organophosphate concentrations, socioeconomic and environmental variables of the observed samples were preprocessed. Missing data were filled using linear interpolation. Outliers were cleaned. All data were logarithmically transformed.
[0059] y = ln(x)
[0060] The data is randomly divided into training and test sets as follows: 70% of all sample data is randomly selected as the training set, and the remaining 30% is the test set. To ensure that the concentration of the training set covers all predicted samples, the three samples with the highest and lowest concentrations are manually added to the training set.
[0061] Step S2: The socioeconomic and environmental variables are screened using the Monte Carlo Uninformation Variable Elimination Method (MC-UVE-RF) based on random forest. Uninformation variables are removed, and variables that have a significant impact on concentration are selected as the model construction dataset.
[0062] In this process, virtual random uninformative variables with five times the number of variables are generated and mixed with real variables; a certain number of samples are randomly drawn from the training set through Monte Carlo simulation as a training subset for building the model, and the posterior inclusion probability of each variable is calculated.
[0063] This invention uses %lncMSE as an indicator to measure the importance of variables, and the importance ranking of variables in a random forest is as follows: Figure 2 As shown, this represents the percentage increase in the mean squared error (MSE) of the model after the variables are randomly shuffled.
[0064]
[0065] The importance index of random forest was used to calculate the posterior inclusion probability in order to screen out socioeconomic and environmental variables that have a significant impact on concentration.
[0066]
[0067] Following the principles of Monte Carlo techniques, most (60%) of the training data should be reserved as a subset of validation data. A cutoff threshold is set to remove variables with no information within the cutoff threshold, while retaining socioeconomic and environmental variables that have a significant impact on concentration. The output is the model building dataset. In this example, the model building dataset includes model building variables such as water residence time, natural flow, temperature, precipitation, altitude, topographic slope, soil clay content, soil organic carbon content, arable land area, and human development index.
[0068] cutoff = k × max[abs(PIP)] noise )]
[0069] Step S3: Construct a random forest concentration quantitative model using the model building dataset and observed sample concentrations. Employ a combination of cross-validation and parameter search, optimize the key parameters of the random forest model based on model evaluation metrics, construct the optimal model, and determine the accuracy of the prediction using model evaluation metrics on the model test set. If the prediction is determined to be inaccurate, rerun MC-UVE-RF and adjust the dummy variable ratio or PIP threshold until the prediction is accurate, determine the model parameters, and obtain a converged organic phosphate concentration quantitative model.
[0070] The prediction index is calculated using a test set. The accuracy of the prediction result is judged based on whether the prediction index meets the index conditions. This invention employs 10-fold cross-validation to calculate the root mean square error (RMSE) or coefficient of determination (R²) of the index. 2 As an estimate of model accuracy, such as Figure 3 As shown, this allows for the selection of optimal model parameters and a determination of the accuracy of model predictions.
[0071]
[0072] Model performance is evaluated using root mean square error (RMSE) and coefficient of determination (CCD). The model with the smallest RMSE and the largest CCD is considered the optimal model. Key parameters such as the number of decision trees and the maximum depth of a single tree are adjusted based on the model evaluation metrics. The criteria for judging prediction accuracy are: the relative RMSE of the optimal model on the test set is less than a preset threshold, and the CCD is greater than a preset threshold.
[0073] Step S4: Input the target area model construction variable dataset into the converged organophosphate concentration quantitative model to predict the organophosphate concentration in the target area.
[0074] Step S5: Use the predicted value of organophosphate concentration to quantify the ecological risk, identify the ecological risk level, and project the risk level of the area with ecological risk onto the target prediction area.
[0075] Among them, the ecological risk is evaluated and classified through the risk quotient (RQ). The risk levels are divided into four levels: no obvious risk (RQ < 0.01), low risk (RQ < 0.1), medium risk (0.1 < RQ < 1), and high risk (RQ > 1).
[0076]
[0077] Example 2:
[0078] An ecological risk prediction system for surface water organophosphate pollution, including:
[0079] Dataset acquisition module: Obtain the dataset of the observed sample organophosphate concentration in the target area and the socio-economic and environmental variable data, preprocess the data, and randomly divide the data into a training set and a test set.
[0080] Variable screening module: Screen the socio-economic and environmental variables through the Monte Carlo uninformative variable elimination method based on random forest (MC-UVE-RF), eliminate the uninformative variables, and screen out the variables that have a significant impact on the concentration as the model construction dataset.
[0081] Model construction module: Use the model construction dataset and the observed sample organophosphate concentration dataset to construct a random forest concentration quantitative model. Adopt a method that combines cross-validation and parameter search to optimize the key parameters of the random forest model according to the model evaluation index, construct an optimal model, and judge whether the prediction is accurate based on the model evaluation index of the model test set. If it is determined that the prediction is inaccurate, re-run MC-UVE-RF and adjust the virtual variable ratio or PIP threshold until the prediction is accurate, determine the model parameters, and obtain the converged organophosphate concentration quantitative model.
[0082] Concentration prediction module: Input the target area model construction variable dataset into the converged organophosphate concentration quantitative model to predict the organophosphate concentration in the target area.
[0083] Environmental ecological risk identification module: Use the predicted value of organophosphate concentration to quantify the ecological risk, identify the ecological risk level, and project the ecological risk level onto the target area.
[0084] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for predicting the ecological risk of surface water organophosphate pollution, characterized in that: Includes the following steps: S1: Obtain the dataset of observed organic phosphate concentrations in the target area, as well as socioeconomic and environmental variable data. Preprocess the data and randomly divide it into training and test sets. S2: Socioeconomic and environmental variables were screened using the Monte Carlo Uninformation Variable Elimination Method (MC-UVE-RF) based on random forest to remove uninformation variables and select variables that have a significant impact on concentration, which were then used as the dataset for model construction. S3: Construct a random forest concentration quantitative model using the model building dataset and the observed sample organic phosphate concentration dataset. Employ a combination of cross-validation and parameter search. Optimize the key parameters of the random forest model based on model evaluation metrics to construct the optimal model. Determine the accuracy of the prediction using model evaluation metrics on the model test set. If the prediction is deemed inaccurate, rerun MC-UVE-RF and adjust the dummy variable ratio or PIP threshold until the prediction is accurate. Determine the model parameters and obtain a converged organic phosphate concentration quantitative model. S4: Input the target region model construction variable dataset into the converged organophosphate concentration quantitative model to predict the organophosphate concentration in the target region; S5: Use predicted values of organophosphate concentrations to quantify ecological risks, identify ecological risk levels, and project the ecological risk levels onto target areas.
2. The method for predicting the ecological risk of surface water organophosphate pollution according to claim 1, characterized in that: S2 specifically includes the following steps: S21: Generate virtual random uninformative variables with 5 times the number of variables, mix them with real variables, and construct a random forest model by randomly sampling a set number of variables using Monte Carlo methods. S22: Calculate the random forest importance index of variables using the random forest model, and calculate the posterior inclusion probability (PIP) of each variable through Monte Carlo simulation based on this, and set the cutoff threshold. S23: Based on the PIP and cutoff threshold, eliminate variables with no information and screen out socioeconomic and environmental variables that have a significant impact on concentration; S24: Output the retained socioeconomic and environmental variables as the model building dataset.
3. The method for predicting the ecological risk of surface water organophosphate pollution according to claim 2, characterized in that: In S22, the random forest importance index for the variables is calculated as follows: by randomly shuffling the values of a variable, the mean squared error of the model is recalculated, and the percentage increase (%lncMSE) of the shuffled mean squared error compared to the original mean squared error is calculated. The index is calculated as follows: Among them, MSE original This is expressed as the root mean square error of the original model. The mean squared error calculated after shuffling the j-th variable is %lncMSE, where %lncMSE represents the percentage increase in MSE after shuffling the eigenvalue, n represents the sample size, and y... i Let y represent the actual concentration of the i-th organophosphate ester. i ' represents the predicted concentration of the i-th organophosphate ester, y i (j) The predicted concentration of organophosphates is obtained by shuffling the j-th variable.
4. The method for predicting the ecological risk of surface water organophosphate pollution according to claim 2, characterized in that: In S22, the posterior inclusion probability and cutoff threshold are calculated as follows: cutoff=k×max[abs(PIP noise )] PIP j Let M represent the posterior inclusion probability of the j-th variable, and M represent the number of Monte Carlo simulations. This represents the random forest importance index for the j-th variable in the m-th sample. I(·) is the indicator function; if the condition within the parentheses is met, I = 1; otherwise, I = 0. It is used to count the effective occurrences. `cutoff` is the cutoff threshold, `k` is the adjustment coefficient, and `PIP` is the key factor. noise This represents the posterior inclusion probability score of an uninformative variable, and abs(·) is the absolute value function.
5. The method for predicting the ecological risk of surface water organophosphate pollution according to claim 2, characterized in that: In step S23, the method for screening socioeconomic and environmental variables that have a significant impact on concentration is to retain the posterior inclusion probability (PIP). j Socioeconomic and environmental variables exceeding the cutoff threshold were used as model building variables, and posterior inclusion probabilities (PIPs) were removed. j Uninformative variables below the cutoff threshold.
6. The method for predicting the ecological risk of surface water organophosphate pollution according to claim 1, characterized in that: The evaluation index in S3 is calculated as follows: The predicted and actual concentrations of organic phosphate esters in the test set are substituted into the formulas for calculating the root mean square error (RMSE) and coefficient of determination to obtain the corresponding RMSE and coefficient of determination. The formulas for calculating the RMSE and coefficient of determination are shown below: relatively Where RMSE represents the root mean square error, R 2 Let y be the coefficient of determination, n be the sample size, and y be the coefficient of determination. i Let y represent the actual concentration of the i-th organophosphate ester. i ′ represents the predicted concentration of the i-th organophosphate ester. This represents the average value of the actual concentration of organic phosphate esters.
7. The method for predicting the ecological risk of surface water organophosphate pollution according to claim 3, characterized in that: The method for determining model parameters in S3 is as follows: construct a parameter search space, select hierarchical k-fold cross-validation based on task type to ensure balanced data distribution; then traverse parameter combinations through grid search or random search, and evaluate model performance using model evaluation metrics; finally, select the optimal parameter combination based on cross-validation results, construct the optimal model, and verify its generalization ability on the test set.
8. The method for predicting the ecological risk of surface water organophosphate pollution according to claim 1, characterized in that: The criterion for judging accurate prediction in S3 is that the relative root mean square error of the optimal model test set is less than a preset threshold and the coefficient of determination is greater than a preset threshold.
9. The method for predicting the ecological risk of surface water organophosphate pollution according to claim 1, characterized in that: In S5, the ecological risk calculation method is to evaluate the ecological risk through the risk quotient (RQ) and classify the levels: where NOEC is the no-observed-effect concentration, LOEC is the lowest-observed-effect concentration, ChV is the chronic no-effect value estimated using the ECOSAR software to ensure consistent environmental risk assessment criteria, RAF is the risk assessment factor, MEC is the predicted environmental concentration, and PNEC is the predicted no-effect concentration; The estimated risk levels are divided into four levels: no obvious risk (RQ < 0.01), low risk (RQ < 0.1), medium risk (0.1 < RQ < 1), and high risk (RQ > 1).
10. A surface water organophosphate pollution ecological risk prediction system, characterized in that: including: Dataset acquisition module: Obtain the dataset of the observed sample organophosphate concentration in the target area and the data of socio-economic and environmental variables, preprocess the data, and randomly divide the data into a training set and a test set; Variable screening module: Screen the socio-economic and environmental variables by the Monte Carlo uninformative variable elimination method based on random forest (MC-UVE-RF), eliminate the uninformative variables, and screen out the variables that have a significant impact on the concentration as the dataset for model construction; Model construction module: Use the dataset for model construction and the dataset of the observed sample organophosphate concentration to construct a random forest concentration quantitative model. Adopt a method combining cross-validation and parameter search to optimize the key parameters of the random forest model according to the model evaluation indexes, construct the optimal model, and judge whether the prediction is accurate according to the model evaluation indexes of the model test set. If the prediction is determined to be inaccurate, re-run MC-UVE-RF and adjust the virtual variable ratio or the PIP threshold until the prediction is accurate, determine the model parameters, and obtain a convergent organophosphate concentration quantitative model; Concentration prediction module: Input the dataset of the model construction variables in the target area into the convergent organophosphate concentration quantitative model to predict the organophosphate concentration in the target area; Environmental ecological risk identification module: Use the predicted value of the organophosphate concentration to quantify the ecological risk, identify the ecological risk level, and project the ecological risk level onto the target area.
Citation Information
Patent Citations
Angelica dahurica decoction piece quality prediction method based on hyperspectral imaging depth analysis
CN113008805A
Organic phosphorus pesticide residue detection method and system based on coupling of electrode method and random forest model, and storage medium
CN117708589A
methanogenesis
US20180237683A1
Methods and Systems for Assessing a Phenotype of a Biological Tissue of a Patient Using Raman Spectroscopy
US20220238187A1