Construction method of prediction model for biomagnification of organic chemicals in low trophic level food chains

Through the construction of the QSAR model, the problem of evaluating the biological amplification effect of organic chemicals in the food chain in the prior art is solved, and the rapid, economical and accurate prediction of biological amplification of PBDEs and HBCDs chemicals is achieved.

CN118800357BActive Publication Date: 2025-06-27NANJING TECH UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410895638.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-05
Publication Date
2025-06-27
Estimated Expiration
2044-07-05

AI Technical Summary

Technical Problem

The prior art is difficult to quickly and economically evaluate the biological amplification effect of organic chemicals in the food chain, especially PBDEs and HBCDs, and the experimental methods are costly and time-consuming.

Method used

Based on the QSAR model, a low-trophic food chain biomagnification prediction model of PBDEs and HBCDs was constructed. By calculating biological enrichment factors (BMF) and other descriptors, a multivariate linear regression model was established to achieve prediction of biological amplification.

Benefits of technology

This model can accurately predict the biological amplification factors of PBDEs and HBCDs chemicals, improve the accuracy of the prediction results, save manpower, material resources and time, and meet the needs of chemical substance ecological risk assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118800357B_ABST
    Figure CN118800357B_ABST
Patent Text Reader

Abstract

The present invention discloses a low-trophic-level food chain biomagnification prediction model for organic chemicals, which is the first low-trophic-level food chain biomagnification prediction model for PBDEs and HBCDs constructed based on the QSAR model. The model is obtained through steps such as sample collection and screening, molecular descriptor calculation, model construction, and model verification. The low-trophic-level food chain biomagnification prediction model for organic chemicals constructed by this method can accurately predict the biomagnification factors of organic pollutants of PBDEs and HBCDs chemicals, improve the accuracy of the prediction results, save manpower, material resources and time, is simple, fast and effective, and strictly follows the QSAR model usage rules stipulated by OECD, explaining the key factors affecting the biomagnification factors from the molecular descriptor structure, which has important significance for the risk control and environmental safety of toxic chemicals such as PBDEs and HBCDs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of environmental risk assessment of organic chemicals, and particularly to a method for constructing a prediction model for biomagnification of organic chemicals in low trophic level food chains. Background Art

[0002] Toxic and harmful chemicals are lipophilic, difficult to degrade, and accumulate and enrich in organisms. In an ecosystem, a biomagnification effect occurs as the trophic level increases, causing toxicity to higher organisms and humans. Therefore, studying the content characteristics of organic pollutants in aquatic organisms is of great significance for the health and food safety of local residents.

[0003] The degree of enrichment of a compound in higher trophic level organisms and humans through the food chain is of great significance for evaluating the ecological and environmental toxicity of the compound. Over the years, scientists have been engaged in related work. Through long-term research, they have discovered and established various compound accumulation models between different media. Generally, there are two criteria for evaluating whether an organic pollutant has a bioaccumulation effect. The first is the K OW (n-octanol-water partition coefficient) of the compound. Generally, when logK OW >4 - 5, the compound may have a bioaccumulation effect, and when logK OW is between 5 - 7, the compound has the maximum bioaccumulation effect; the second is the bioconcentration factor (BAF). The bioconcentration factor (BAF) can characterize the relative bioaccumulability of a compound.

[0004] The calculation formula for the bioconcentration factor BAF in an aquatic food chain is as follows:

[0005]

[0006] C 生物 represents the concentration of the pollutant in the aquatic organism, with the unit pg / kg lw; C 水中溶解相 represents the concentration of the dissolved pollutant in the water body, with the unit pg / L. The concentration in the organism is the lipid-normalized concentration.

[0007] In the BAF model, when the BAF value of a certain pollutant is higher than 5000 (or LogBAF > 3.7), it can be considered that the pollutant has a bioaccumulation effect in the food chain; when the BAF value is between 2000 - 5000 (or LogBAF > 3.7), it is considered to have a potential bioaccumulation effect.

[0008] As newly listed new persistent organic pollutants under the Stockholm Convention, PBDEs and HBCDs have typical bioaccumulation properties. Due to the characteristics of high lipophilicity and poor metabolic ability of PBDEs and HBCDs, they have the potential to magnify along the food chain. Recent studies have also confirmed that PBDEs and HBCDs can be transferred along the food chain with trophic levels. However, obtaining the bioconcentration magnification factor (BMF) of chemicals only through experimental methods is costly, time-consuming and laborious, and it is difficult to meet the needs of ecological risk assessment of chemical substances. At present, the model for the biomagnification effect of chemical substances in the food chain is still vacant. Therefore, there is an urgent need to develop a scientific, fast and effective theoretical calculation method for the food chain transfer model. The Organization for Economic Co-operation and Development (OECD) issued guidelines for the construction and validation of QSAR models in 2007, and put forward the criteria that QSAR models should meet: ① having a clearly defined environmental indicator; ② having a clear and definite mathematical algorithm; ③ defining the application domain of the model; ④ the model having appropriate goodness of fit, robustness and predictive ability; ⑤ making model mechanism explanations as much as possible. Summary of the Invention

[0009] For the first time based on the QSAR model, the present invention constructs a biomagnification prediction model for low trophic level food chains of PBDEs and HBCDs, as follows:

[0010] PEC oral,predator = PEC water *BCF fish *BMF (1)

[0011] BMF=-5.04472 + 0.8374*GGI3 - 35.46426*Mor21v (2)

[0012] The meanings of the parameters are as follows:

[0013]

[0014] The present invention provides a method for constructing the above-mentioned biomagnification model for low trophic level food chains, specifically as follows:

[0015] ⑴ Sample collection and screening

[0016] The laboratory data obtained includes the biomagnification factor data of 9 organic compounds, and these compounds cover the organic matters of PBDEs and HBCDs. In order to establish an effective QSAR model, the data set is first divided into a training set and a test set. To ensure the representativeness of the compounds in the training set, the grouping method used in this work is the Kennard&Stone method. This method can avoid the uneven distribution of training set samples to a certain extent and can well divide the data set into a training set and a test set. This data set is divided into 6 training sets and 3 validation sets.

[0017] ⑵ Molecular descriptor calculation

[0018] In this method, the molecular structures of 9 organic compounds were first constructed in ChemDraw software and then imported into the HyperChem program for molecular optimization. The optimization was divided into two steps: first, the MM+ molecular force field method was used for preliminary energy optimization, and then the semi-empirical quantum mechanics AM1 method was used for more accurate conformational optimization of the structure. The optimized structure was imported into the DRAGON 5.4 software to calculate 1664 different types of theoretical molecular descriptors. These descriptors were preprocessed before modeling, that is, the constant terms, terms close to constants, and molecular descriptors with high correlation (the one with a smaller correlation coefficient with the target value among two molecular descriptors with a correlation coefficient greater than 0.96) were deleted. Finally, 1169 descriptors remained for the subsequent variable selection process.

[0019] ⑶ Model construction

[0020] The genetic algorithm was used to select the descriptor set highly correlated with bioaccumulation, and this process was implemented in MobyDigs. After variable selection by the genetic algorithm, a linear QSAR model, namely the GA-MLR model, was established using the multiple linear regression (MLR) method. The model evaluation function was selected as the leave-one-out cross-validation, that is, when the performance of the model did not change significantly after adding a descriptor (the increase in Q2 was less than 0.02 when adding a descriptor), the optimal number of descriptors was reached. In this method, the optimal number of descriptors was 7. The relevant parameters in modeling were set as follows: the population size was 100, the maximum allowed variables for the initial model was 7, the mutation trade-off (T) was 0.5, and the crossover and mutation probabilities were both based on the T parameter.

[0021] (4) Model validation

[0022] After variable selection by the genetic algorithm, a linear QSAR model, namely the MLR model, was established using the multiple linear regression method. The linear MLR equation is as follows:

[0023] BMF = -5.04472 + 0.8374 * GGI3 - 35.46426 * Mor21v

[0024] n tr = 6Q 2 LOO = 0.9981R 2 fitting= 0.9996R 2 adj = 0.9994RMSE tr = 0.0665R 2 boot = 0.7175

[0025] n ext = 3R 2 ext = 0.836, Q 2 ext = 0.8662R 2 adj = 0.9994RMSE ext = 0.0289

[0026] Among them, GGI3 represents the 3rd-order topological charge index, Mor21v represents the 3D-MoRSE-weighted atomic van der Waals volume, GGI3 has a positive correlation with the biomagnification factor, Mor21v has a negative correlation with the biomagnification factor, n tr and n ext are the numbers of compounds in the training set and the validation set respectively; R 2 adj is the coefficient of determination corrected for degrees of freedom; RMSE tr is the root mean square error of the training set; Q 2 LOO is the leave-one-out cross-validation coefficient; R 2 boot is the validation coefficient by the Bootstrapping method; R 2 ext is the coefficient of determination between the experimental values and the predicted values in the validation set, Q 2 ext is the external validation coefficient of determination, RMSE ext is the root mean square error of the validation set, SE is the standard error. R 2 fitting is the coefficient of determination between the experimental values and the predicted values in the training set. The RMSEs of the training set and the validation set are 0.0665 and 0.0289 respectively, and the model prediction effect is good.

[0027] Table 1 Experimental values and predicted values of BMFs of PBDEs and HBCDs

[0028] Substance CAS Training set / Validation set Experimental value Predicted value BDE100 189084-64-8 Training set 4.07 4.07 BDE153 68631-49-2 Validation set 1.14 1.81 BDE154 207122-15-4 Training set 3.05 3.05 BDE183 207122-16-5 Validation set 1.5 1.33 BDE28 41318-75-6 Training set 3.6 3.65 BDE47 5436-43-1 Training set 3.46 3.42 BDE99 60348-60-9 Validation set 3.85 4.53 BPE209 145538-74-5 Training set 6.04 6.05 HBCD 3194-55-6 Training set 7.26 7.25

[0029] The advantages of the low trophic level food chain biomagnification prediction model for organic chemicals established by this method are as follows: By experimental means, the concentration of pollutants PBDEs and HBCDs in aquatic organisms and the concentration of dissolved pollutants in water are measured, and then the biomagnification time of chemicals in the food chain is calculated, which is time-consuming and costly. The low trophic level food chain biomagnification prediction model for organic chemicals constructed by this method can accurately predict the biomagnification factors of organic pollutants such as PBDEs and HBCDs, improve the accuracy of prediction results, save manpower, material resources and time, is simple, fast and effective, and strictly follows the QSAR model usage rules stipulated by OECD. It explains the key factors affecting the biomagnification factor from the molecular descriptor structure, which is of great significance for the risk control of toxic chemicals such as PBDEs and HBCDs and environmental safety. Description of the Drawings

[0030] Figure 1 Fitting diagram of the BMF prediction model.

[0031] Figure 2 Characterization diagram of the application domain of the BMF prediction model. Detailed Implementation Modes

[0032] The present invention will be further described below in conjunction with specific embodiments.

[0033] Example 1

[0034] The specific steps for constructing the low trophic level food chain biomagnification prediction model for organic chemicals are as follows:

[0035] (1) Data collection, setting training set and validation set sample compounds

[0036] The biomagnification factors BMF of 9 organic chemicals were obtained in the laboratory. A total of 6 sample compounds were selected for the training set, and 3 sample compounds were selected for the validation set.

[0037] (2) Calculating descriptors

[0038] The MM+ molecular mechanics in Hyperchem 7.0 software was used to pre-optimize the compound structure, and the semi-empirical AM1 method was used to optimize the compound structure. Based on the optimized structure, Dragon 5.4 software was used to calculate descriptors, and 1664 calculated descriptors were preliminarily screened.

[0039] (3) Model construction

[0040] The genetic algorithm (GA) in MobyDigs software was used for variable selection. Based on the selected variables, the multiple linear regression (MLR) method was used to establish a prediction model, namely the GA-MLR model:

[0041] BMF = -5.04472 + 0.8374 * GGI3 - 35.46426 * Mor21v

[0042] Among them, GGI3 represents the third-order topological charge index, Mor21v represents the 3D-MoRSE-weighted atomic van der Waals volume. GGI3 has a positive correlation with the biomagnification factor, and Mor21v has a negative correlation with the biomagnification factor. The RMSE values of the training set and the validation set are 0.0665 and 0.0289 respectively.

[0043] (4) Model Validation

[0044] According to the OECD guidelines for QSAR models, the constructed model needs to be internally validated (assessment of goodness of fit and robustness) and externally validated (assessment of predictive ability). The square of the correlation coefficient (R 2 adj ) between the corrected experimental values and the fitted values, and the root mean square error (RMSE) are used to characterize the goodness of fit of the model:

[0045]

[0046] Among them, n represents the number of compounds, m is the number of predictive variables, y i and represent the experimental value and the predicted value of the activity index of the i-th compound respectively; is the average value of the experimental values of the compound activity index.

[0047] The leave-one-out cross-validation coefficient (Q 2 LOO ) and the Bootstrapping method (Q 2 BOOT ) are used to characterize the stability of the model:

[0048]

[0049] Among them, represents the average value of the experimental values of the training set compound activity index. The Bootstrapping method uses 1 / 5 leave-one-out cross-validation and repeats 5000 times.

[0050] The external validation correlation coefficient (Q 2 EXT ), R 2 EXT , RMSE EXT are used to characterize the predictive ability of the model:

[0051]

[0052] Among them, n EXTRepresents the number of compounds in the validation set, Indicates the average of the experimental values and predicted values of the activity indicators of the compounds in the validation set. The characterization and evaluation parameters of the model are obtained:

[0053] n tr = 6Q 2 LOO = 0.9981R 2 fitting = 0.9996R 2 adj = 0.9994RMSE tr = 0.0665R 2 boot = 0.7175

[0054] n ext = 3R 2 ext = 0.836, Q 2 ext = 0.8662R 2 adj = 0.9994RMSE ext = 0.0289

[0055] Among them, n tr and n ext are the numbers of compounds in the training set and validation set respectively, and p is the significance level. R 2 adj is the coefficient of determination corrected for degrees of freedom; RMSE is the root mean square error; Q 2 LOO is the leave-one-out cross-validation coefficient; Q 2 BOOT is the validation coefficient by the Bootstrapping method; R 2 ext is the correlation coefficient between experimental values and predicted values, Q 2 ext is the external validation coefficient of determination, and RMSE ext is the root mean square error of the validation set.

[0056] The results show that the model has good predictive ability and robustness.

[0057] Example 2

[0058] In this example, the application domain of the above prediction model is characterized.

[0059] The Williams plot is a model application domain defined by the standardized residuals (δ) and leverage values (denoted by h i , where i represents different compounds). δ is calculated by the following formula:

[0060]

[0061] The leverage value (h i ) of the training set compounds can be obtained by the following formula:

[0062] h i = x i T (X T X) –1 x i (8)

[0063] In the formula, x i is the row vector of the molecular structure descriptors of the i-th compound. The warning value (h * ) is defined as:

[0064] h * = 3(k + 1) / n (9)

[0065] where k is the number of descriptors and n is the number of training sets.

[0066] The results of characterizing the model application domain are as Figure 1 、 Figure 2 shown. Figure 1 In * , h i =3(k+1) / n=3(2+1) / 6=1.5. The ordinate of the Williams plot characterizes the dispersion degree of the experimental values with the standard residuals of the experimental values and the predicted values. When the absolute value of the standard residual δ of the compound is greater than 3.0, it is regarded as an outlier. The abscissa represents the h i value of the compounds in the training set. When h

[0067] is greater than the warning value (h* = 1.5), it indicates that the substructure of this substance appears less in the training set and will have a significant impact on the model prediction results. * As can be seen from the figure, the leverage value h of all compounds is within the warning leverage value h Figure 2 . This indicates that the structure of this compound has a certain similarity to the structures of the training set compounds. The standard residuals are all within the range of (-3, +3), indicating that this model is applicable to the prediction of BDE153 (CAS: 68631-49-2), BDE 183 (CAS: 207122-16-5) and BDE-99 (CAS: 60348-60-9). The standard residuals of all fall within the range of (-3, +3), indicating that this model is applicable to the prediction of these three substances and can be well predicted. See

[0068]

[0069] Example 3

[0070] The model constructed in Example 1 was used to predict the food chain magnification factors of 9 persistent organic pollutants, and the results are shown in Table 2. The R 2 adj of the model = 0.999, indicating that the model has a strong fitting ability. Q 2 LOO = 0.9981, Q 2 BOOT = 0.7175, indicating that the model is relatively robust. R 2 ext = 0.9994, Q 2 ext = 0.8662. Golbraikh et al. believe that the acceptable criteria for QSAR models are Q 2 > 0.50 and R 2 > 0.60. The results show that the model has good predictive ability and can be successfully applied to compounds outside the training set. Figure 1 is a fitting graph of the predicted value and experimental value of the food chain magnification factor BMF of persistent organic chemicals. As can be seen from Figure 1 it, the predicted values and experimental values of most substances fit well, and BDE153 (CAS: 68631-49-2), BDE183 (CAS: 207122-16-5) and BDE-99 (CAS: 60348-60-9) can be predicted well.

[0071] Example 4

[0072] Using the model constructed in Example 1, the biomagnification factor BMF of BDE 203 (SMILES: BrC1=C(OC2=CC(Br)=C(Br)C(Br)=C2Br)C(Br)=CC(Br)=C1Br) was predicted. First, according to the molecular structure of the chemical substance, two descriptors GGI3 and Mor21v were calculated using Dragon software, which were 1.937 and -0.06 respectively, and Hat was 0.435, within the scope of the model application domain.

[0073] BMF = -5.04472 + 0.8374 * GGI3 - 35.46426 * Mor21v

[0074] BMF = -5.04472 + 0.8374 * (1.937) - 35.46426 * (-0.06) = 3.25

[0075] Then the predicted value of BDE 203 (BMF) is 3.25, which is close to the test measurement result (1.73).

[0076] Example 5

[0077] Using the model constructed in Example 1, predict the biomagnification factor BMF of BDE 196 (SMILES: BrC1=CC=C(OC2=C(Br)C(Br)=C(Br)C(Br)=C2Br)C(Br)=C1Br). First, according to the molecular structure of the chemical substance, two descriptors GGI3 and Mor21v were calculated using Dragon software, which were 2.375 and -0.016 respectively, and Hat was 0.651, within the scope of the model application domain.

[0078] BMF = -5.04472 + 0.8374 * GGI3 - 35.46426 * Mor21v

[0079] BMF = -5.04472 + 0.8374 * (2.375) - 35.46426 * (-0.016) = 2.05

[0080] The predicted value of BDE 196 (BMF) is 2.05, which is close to the experimental measurement result (1.43).

Claims

1. A method for constructing a low trophic level food chain biomagnification prediction model for organic chemicals, characterized in that: The low trophic level food chain biomagnification prediction model is as follows: PEC oral,predator = PEC water *BCF fish *BMF (1) BMF=-5.04472+0.8374*GGI3-35.46426*Mor21v (2) Among them, PEC oral,predator Refers to the concentration in the body of low trophic level predators, mg kg wetfish -1 ;PEC water Refers to the predicted concentration in water, mg / L; BCF fish Refers to the bioconcentration factor of fish, L·kg wetfish -1 ; BMF refers to biological magnification factor; GGI3 refers to 3rd-order topological charge index; Mor21v refers to 3D-MoRSE-weighted atomic van der Waals volume; The method for constructing the low trophic level food chain biomagnification prediction model of organic chemicals comprises the following steps: (1) Sample collection and screening; (2) Calculation of molecular descriptors; (3) Model construction; (4) Model validation; Step (1) is as follows: laboratory data are obtained containing biomagnification factor data of 9 organic compounds, which include organic compounds of PBDEs and HBCDs. First, the data set is divided into a training set and a test set. The grouping method used is the Kennard & Stone method, which is divided into 6 training sets and 3 validation sets; Step (2) is as follows: first, the molecular structures of 9 organic compounds are constructed in ChemDraw software, and then the molecules are optimized by importing the HyperChem program; the optimization is divided into two steps: first, the MM+ molecular force field method is used for preliminary energy optimization, and then the semi-empirical quantum mechanics AM1 method is used to optimize the structure more accurately, and the optimized structure is imported into DRAGON5.4 software to calculate 1664 different types of theoretical molecular descriptors; these descriptors are preprocessed before modeling, that is, constant terms, terms close to constants, and highly correlated molecular descriptors are deleted, and finally the remaining 1169 descriptors are used for the subsequent variable selection process; Step (3) is as follows: a genetic algorithm is used to select a set of descriptors that are highly correlated with bioaccumulation. This process is implemented in MobyDigs. After the genetic algorithm variable selection, a linear QSAR model is established using the multivariate linear regression MLR method. The model evaluation function selects the leave-one-out interaction test, that is, when the performance of the model does not change significantly after adding a descriptor, the optimal number of descriptors is reached. In this method, the optimal number of descriptors is 7. The relevant parameters in the modeling are set as follows: population size is 100, the maximum allowed variables allowed in the initial model is 7, the mutation trade-off value T is 0.5, and the crossover and mutation probabilities are both based on the T parameter. Step (4) is as follows: after genetic algorithm variable selection, a linear QSAR model, namely, an MLR model, is established using the multivariate linear regression method. The linear MLR equation is as follows: BMF=-5.04472+0.8374*GGI3-35.46426*Mor21v <h2 style=";text-align:left;direction:ltr">n<h2 style=";text-align:left;direction:ltr"> tr <h2 style=";text-align:left;direction:ltr"> <6Q<h2 style=";text-align:left;direction:ltr"> 2 <h2 style=";text-align:left;direction:ltr"> LOO <h2 style=";text-align:left;direction:ltr"> <0.9981R<h2 style=";text-align:left;direction:ltr"> 2 <h2 style=";text-align:left;direction:ltr"> fitting <h2 style=";text-align:left;direction:ltr"> <0.9996R<h2 style=";text-align:left;direction:ltr"> 2 <h2 style=";text-align:left;direction:ltr"> adj <h2 style=";text-align:left;direction:ltr"> <0.9994RMSE<h2 style=";text-align:left;direction:ltr"> tr <h2 style=";text-align:left;direction:ltr"> <0.0665R<h2 style=";text-align:left;direction:ltr"> 2 <h2 style=";text-align:left;direction:ltr"> boot <h2 style=";text-align:left;direction:ltr"> <0.7175 <h2 style=";text-align:left;direction:ltr">n<h2 style=";text-align:left;direction:ltr"> ext <h2 style=";text-align:left;direction:ltr"> <3R<h2 style=";text-align:left;direction:ltr"> 2 <h2 style=";text-align:left;direction:ltr"> ext <h2 style=";text-align:left;direction:ltr"> =0.836,Q<h2 style=";text-align:left;direction:ltr"> 2 <h2 style=";text-align:left;direction:ltr"> ext <h2 style=";text-align:left;direction:ltr"> =0.8662R<h2 style=";text-align:left;direction:ltr"> 2 <h2 style=";text-align:left;direction:ltr"> adj <h2 style=";text-align:left;direction:ltr"> <0.9994RMSE<h2 style=";text-align:left;direction:ltr"> ext <h2 style=";text-align:left;direction:ltr"> <0.0289 Among them, GGI3 represents the third-order topological charge index, Mor21v represents the 3D-MoRSE-weighted atomic van der Waals volume, GGI3 is positively correlated with the biomagnification factor, Mor21v is negatively correlated with the biomagnification factor, and the RMSE of the training set and validation set are 0.0665 and 0.0289, respectively; n tr and n ext are the number of compounds in the training set and the validation set, respectively; R 2 adj is the coefficient of determination corrected for degrees of freedom; RMSE tr is the root mean square error of the training set; Q 2 LOO is the leave-one-out cross-validation coefficient; R 2 boot is the validation coefficient of the Bootstrapping method; R 2 ext is the coefficient of determination between the experimental value and the predicted value of the validation set, Q 2 ext is the external validation coefficient of determination, RMSE ext is the root mean square error of the validation set, SE is the standard error, R 2 fitting is the coefficient of determination between the experimental value and the predicted value of the training set.

Citation Information

Patent Citations

  • Method for predicting fish bio-concentration factors of organic chemicals by quantitative structure-activity relationship

    CN103761431A

  • Method for directly predicting biological effectiveness of organic pollutants

    CN110619925A