Method for predicting toxicity of commercial pesticide to earthworms based on machine learning

By building a commercial pesticide toxicity prediction model based on machine learning and integrating the characteristics of commercial pesticide technical, adjuvants and soil properties, the problems of high time consumption and high cost of traditional evaluation methods are solved, rapid and accurate pesticide toxicity assessment is achieved, and scientific environmental risk management support is provided.

CN120656597APending Publication Date: 2025-09-16NANJING INST OF ENVIRONMENTAL SCI MINIST OF ECOLOGY & ENVIRONMENT OF THE PEOPLES REPUBLIC OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510678692.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Traditional commercial pesticide toxicity assessment methods are time-consuming and costly, making it difficult to conduct high-throughput screening of thousands of commercial pesticides and their compound adjuvants. Existing risk assessments also ignore the potential effects of adjuvants on non-target organisms.

Method used

A machine learning-based approach was used to construct a toxicity prediction model. This model integrated the molecular characterization features of commercial pesticide technicals, adjuvant types, and soil physical and chemical properties. Key features were screened using an extreme gradient boosting model. Model training and Bayesian optimization were combined with multiple machine learning algorithms to construct an efficient toxicity prediction model.

Benefits of technology

It has achieved rapid and accurate prediction of the toxicity of commercial pesticides under actual soil conditions, reduced reliance on experimental testing, improved prediction accuracy, revealed the key factors affecting the toxicity of commercial pesticides, and provided a scientific basis for the safe use of pesticides and environmental management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656597A_ABST
    Figure CN120656597A_ABST
Patent Text Reader

Abstract

The invention discloses a method for predicting toxicity of commercial pesticides to earthworms based on machine learning, which comprises the following steps: acquiring a data set, and constructing a toxicity prediction model based on the data set; the data set comprises the toxicity data of the commercial pesticide to the earthworms, the molecular characterization characteristics of the original pesticide of the commercial pesticide, the assistant type of the commercial pesticide and the physical and chemical properties of soil; based on the toxicity prediction model, obtaining the predicted toxicity level of the commercial pesticide to the earthworms; the method provided by the invention makes up for the defect of neglecting the types of auxiliaries in the prior art, can efficiently and accurately predict the toxicity of commercial pesticides to non-target organisms under actual soil conditions, and provides a scientific basis for safe use of pesticides and environmental risk management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of commercial pesticide hazard prediction, and in particular to a method for predicting the toxicity of commercial pesticides to earthworms based on machine learning. Background Art

[0002] Commercial pesticides have become key chemicals in modern agricultural production due to their high effectiveness in suppressing pests and diseases. Appropriate application can significantly increase grain yields, ensure agricultural product quality, and promote economic development. However, improper or excessive use can severely damage soil ecosystems. Earthworms, as important functional organisms in the soil, not only participate in the decomposition of organic matter and nutrient cycling but also influence soil structure and aeration. Their health directly reflects the quality of the soil environment.

[0003] Traditional toxicity assessments rely primarily on laboratory acute or chronic exposure tests. These methods are time-consuming and costly, making them unsuitable for high-throughput screening of the tens of thousands of commercial pesticides and their adjuvant combinations on the market. Furthermore, existing risk assessments typically focus solely on the active ingredient of the technical pesticide, ignoring the potential effects of adjuvants on non-target organisms. Therefore, there is an urgent need to develop a rapid and low-cost predictive technology that can account for both technical pesticide and adjuvant types to support the safety evaluation and precise regulation of commercial pesticide formulations. Summary of the Invention

[0004] In order to solve the above problems, the present invention provides a method for predicting the toxicity of commercial pesticides to earthworms based on machine learning.

[0005] A method for predicting the toxicity of commercial pesticides to earthworms based on machine learning, comprising the following steps:

[0006] Acquire a data set; the data set includes: toxicity data of commercial pesticides to earthworms, molecular characterization characteristics of commercial pesticide technicals, types of commercial pesticide adjuvants, and soil physical and chemical properties;

[0007] Constructing a toxicity prediction model based on the data set; the input of the toxicity prediction model is the molecular characterization characteristics of the commercial pesticide technical, the type of commercial pesticide adjuvant, and the physical and chemical properties of the soil; the output of the toxicity prediction model is whether the commercial pesticide is toxic to earthworms;

[0008] Based on the toxicity prediction model, predict whether commercial pesticides are toxic to earthworms.

[0009] Description: This method efficiently and accurately predicts the toxicity of commercial pesticides under realistic soil conditions, providing a scientific basis for their safe use and environmental risk management. By integrating multi-source data on commercial pesticide adjuvant type, technical properties, and soil properties as inputs for model prediction, this method not only improves prediction accuracy but also reveals key factors influencing commercial pesticide toxicity, providing guidance for their rational design and use. Furthermore, this method reduces reliance on expensive and time-consuming experimental testing, accelerates the process of commercial pesticide toxicity assessment, and facilitates the timely identification and prevention of potential environmental risks.

[0010] Furthermore, the toxicity data of the commercial pesticide to earthworms is the concentration of the commercial pesticide in the unit soil that causes 50% of the earthworms to die within 14 days LC 50 value.

[0011] Description: By using the above LC 50 The values ​​of these values ​​can provide a more accurate understanding of the potential harm of commercial pesticides to earthworms, thereby formulating more effective environmental protection measures and commercial pesticide management policies.

[0012] Furthermore, commercial pesticides include: herbicides, insecticides, fungicides, acaricides, nematicides, and plant growth regulators; the adjuvant types of commercial pesticides include dispersible water granules, water emulsions, suspensions, water-dispersible powders, aqueous solutions, smoke sprays, microcapsule powders, microcapsule suspensions, oil dispersants, effervescent tablets, seed dressings, water-soluble granules, soluble concentrates, soluble powders, concentrates, granules, wettable powders, microemulsions, and emulsifiable concentrates.

[0013] Note: The above content lists the adjuvant types of commercial pesticides in detail, providing comprehensive input features for building toxicity prediction models, so that the models can more accurately capture the chemical properties and potential ecological risks of commercial pesticides, and help improve the predictive ability of the models.

[0014] Furthermore, the soil properties include soil organic matter and soil pH.

[0015] Note: The above settings identify key factors in soil properties that have a significant impact on the behavior of commercial pesticides in soil and their toxicity to earthworms.

[0016] Furthermore, the data set is obtained through experiments or data query, and the toxicity data of commercial pesticides to earthworms in the data set include LC 50 Non-toxic data with a value of ≥100 mg / kg dry soil, LC values ​​of commercial pesticides for earthworms 50 Toxicity data with values ​​< 100 mg / kg dry soil.

[0017] Description: By clarifying thresholds, complex toxicity data can be simplified and standardized labels can be provided for machine learning models to achieve efficient risk assessment and decision support.

[0018] Furthermore, the molecular characterization features include molecular descriptors and molecular fingerprints; the molecular descriptors include Mordred descriptors and RDKit descriptors, and the molecular fingerprints include Morgan fingerprints.

[0019] Note: The above three molecular characterization features can form a complementary feature space, avoiding the limitations of a single method and improving the efficiency and reliability of molecular modeling.

[0020] Furthermore, the method for constructing a toxicity prediction model based on the data set includes:

[0021] Mark the toxic data in the data set as the toxic group and the non-toxic data as the non-toxic group;

[0022] The extreme gradient boosting model was used to screen key molecular features from the molecular features of commercial pesticide technicals. Correlation analysis was then performed to eliminate key molecular features with high autocorrelation, resulting in the screening of molecular features and the corresponding commercial pesticide adjuvant types, soil physical and chemical properties, and toxicity data of the commercial pesticides to earthworms.

[0023] Model training is performed using M machine learning algorithms on the molecular characterization features after screening the three conditions, as well as the corresponding commercial pesticide adjuvant types, soil physical and chemical properties, and commercial pesticide toxicity data to earthworms, to construct M×3 candidate prediction models;

[0024] Based on the toxic group and the non-toxic group, ten-fold cross validation was used for internal validation, and then Bayesian optimization was used to tune each of the candidate prediction models; the accuracy, recall rate and F1 score indicators of the candidate prediction models were calculated, and the optimal model was selected as the toxicity prediction model based on the accuracy, recall rate and F1 score indicators.

[0025] Description: The above method can effectively improve the predictive accuracy and reliability of the model by using multiple molecular characterization feature data, extreme gradient boosting models to screen key molecular descriptors / molecular fingerprints, correlation analysis to eliminate redundant information, and Bayesian optimization to tune the model. Through multi-step, multi-level data processing and model optimization strategies, it not only ensures that the model can capture the complex factors affecting the toxic effects of commercial pesticides, but also improves the model's generalization ability, providing a powerful tool for environmental risk assessment of commercial pesticides.

[0026] Furthermore, the method of selecting the optimal model as the toxicity prediction model based on the accuracy, recall and F1 score indicators includes:

[0027] First, according to the accuracy ranking, at least 6 candidate prediction models with higher accuracy are selected from M×3 candidate prediction models;

[0028] Among the at least six candidate prediction models with relatively high accuracy, a candidate prediction model with the largest sum of recall rate and F1 score is selected as the toxicity prediction model.

[0029] Explanation: Initial accuracy screening ensures high overall model prediction accuracy. Screening by the sum of recall and F1 scores improves coverage of toxic samples and the balance between precision and recall, avoiding missed high-risk samples due to bias in a single indicator. This is because misclassifying a non-toxic pesticide as toxic can lead to unnecessary toxicity testing, while misclassifying a toxic pesticide as non-toxic can pose significant environmental risks. Therefore, the present invention can more accurately predict commercial pesticides that are toxic to earthworms through this judgment.

[0030] Furthermore, before making a prediction based on the toxicity prediction model, the reliability of the prediction is first determined. Methods for determining the reliability of the prediction include:

[0031] First, determine the applicable domain threshold of the molecular characterization features of commercial pesticide technicals in the toxicity prediction model;

[0032] Calculate the average Euclidean distance between the molecular characteristics of the commercial pesticide technical in the input of the toxicity prediction model and the molecular characteristics of the commercial pesticide technical in the nearest neighbor sample in the training set;

[0033] When the average Euclidean distance is less than or equal to the applicable domain threshold, the prediction result is considered reliable; conversely, when the average Euclidean distance is greater than the applicable domain threshold, the prediction result is considered unreliable.

[0034] Note: The above method determines the applicability threshold and calculates the average Euclidean distance between the input molecule and its nearest neighbor in the training set. This method effectively determines whether the model's predictions on new data are within its training range. This reliability assessment helps avoid inaccurate predictions outside the model's training range, thereby improving the model's credibility and security in real-world applications.

[0035] Furthermore, the calculation formula of the applicable domain threshold is as follows:

[0036]

[0037] In formula (1), is the average Euclidean distance between the molecular characterization feature of each commercial pesticide technical in the training set of the dataset and its three nearest neighbors, σ is the standard deviation of the Euclidean distances of the three nearest neighbors, and Z is a constant with a value range of 0.4 to 0.6.

[0038] Note: Calculating the domain threshold using the mathematical formula above provides a quantitative criterion for evaluating the reliability of toxicity prediction models. This formula can determine a reasonable threshold for determining whether new data falls within the model's domain. This quantitative approach ensures that the model's predictions are more reliable and trustworthy in practical applications.

[0039] Furthermore, the calculation formula of the prediction accuracy is as follows:

[0040]

[0041] Where TP represents the number of data that are actually toxic and are also predicted to be toxic by the toxicity prediction model; TN represents the number of data that are actually non-toxic and are also predicted to be non-toxic by the toxicity prediction model; FP represents the number of data that are actually non-toxic but are predicted to be toxic by the toxicity prediction model; and FN represents the number of data that are actually toxic but are predicted to be non-toxic by the toxicity prediction model.

[0042] Note: The above method defines the calculation formula for prediction accuracy, allowing researchers and users to accurately evaluate and compare the performance of different toxicity prediction models.

[0043] The beneficial effects of the present invention are:

[0044] The method presented in this paper can efficiently and accurately predict the toxicity of commercial pesticides under actual soil conditions, thus providing a scientific basis for their safe use and environmental risk management. By integrating multi-source data, such as the type of commercial pesticide adjuvant, the properties of the technical pesticide, and soil properties, into the model prediction, the accuracy of the predictions is improved and the key factors influencing the toxicity of commercial pesticides are revealed, providing guidance for their rational design and use. Furthermore, this method reduces reliance on expensive and time-consuming experimental testing, accelerates the process of commercial pesticide toxicity assessment, and facilitates the timely identification and prevention of potential environmental risks. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 This is a flow chart of prediction based on a toxicity prediction model according to an embodiment of the present invention;

[0046] Figure 2 is the accuracy of the 30 models in the embodiment of the present invention on the training set. (a) (b) (c) are the accuracy of the classifiers built using Mordred descriptor, RDKit descriptor and Morgan fingerprint on the training set respectively;

[0047] Figure 3 The figure shows the accuracy of the two-category toxicity prediction models in the embodiment of the present invention on the test set. The outer circle shows the accuracy of 30 models, and the box plot in the inner circle shows the distribution of the accuracy when each machine learning framework is combined with the three molecular characterization features.

[0048] Figure 4 The confusion matrices of the six models with relatively high accuracy in the embodiment of the present invention are shown in Figure 2. (a) and (d) represent the Bagging and RF models based on the Mordred descriptor, respectively. (b), (c), (e), and (f) represent the Bagging, RF, GBDT, and LightGBM models using the RDKit descriptor.

[0049] Figure 5 Figure 2 is the ROC curve of six models with relatively high accuracy in the embodiment of the present invention; (a) represents bagging and the RF model based on Mordred descriptor; (b) represents bagging; RFGBDT and LightGBM models using RDKit descriptor, and AUROC values ​​are provided to quantify the performance of these models in (a) and (b);

[0050] Figure 6 This is an analysis of the applicable domain and interpretability of the GBDT model constructed using the RDKit descriptor in an embodiment of the present invention. DETAILED DESCRIPTION

[0051] In order to further illustrate the approach and effects achieved by the present invention, the technical solution of the present invention will be clearly and completely described below in conjunction with experiments.

[0052] Commercial pesticides play a significant role in agricultural production but may pose significant risks to non-target soil organisms, such as earthworms. Traditional experimental methods are time-consuming and costly, making it difficult to evaluate the large number of commercial pesticides available on the market. Although quantitative structure-activity relationship (QSAR) prediction models have been established, existing models and methods rarely consider the composition of commercial pesticides and their toxicity under natural soil conditions. For example, existing QSAR models, which are important tools for predicting the activity and properties of chemicals, typically rely on two molecular characterization methods, molecular descriptors (MDs) and molecular fingerprints (MFs). (Molecular characterization features in this article refer to molecular characterization methods such as molecular descriptors (MDs) and molecular fingerprints (MFs), to predict the toxicity of chemicals. Commercial pesticides are often formulated with additives such as surfactants, stabilizers, emulsifiers, dispersants, and other chemicals to enhance their efficacy. These adjuvants may also influence the toxicity of commercial pesticides to soil organisms, but their impact is often overlooked in research and regulatory assessments. Furthermore, traditional QSAR models are limited to the molecular properties of chemicals, making it difficult to more reliably predict the toxicity of commercial pesticides to non-target organisms in actual soil environments. Therefore, when evaluating the environmental toxicity of commercial pesticides in the embodiments of the present invention, the molecular characteristics of the commercial pesticide technical, environmental factors, and adjuvants of the commercial pesticide are taken into consideration.

[0053] Earthworms are essential components of soil ecosystems, acting as ecosystem engineers involved in maintaining soil structure, decomposing organic matter, cycling nutrients, and regulating water. They can survive in nearly all soil types and are widely considered sensitive indicators of soil disturbance.

[0054] Therefore, in the present embodiment, Eisenia fetida was selected as a non-target soil organism of commercial pesticides, and a machine learning-based commercial pesticide toxicity prediction model (ML-QSAR) was developed. By incorporating the molecular characteristics of commercial pesticides, soil properties, and the types of commercial pesticide adjuvants, the toxicity of commercial pesticides to earthworms under actual soil conditions was predicted.

[0055] Combining the above content, such as Figure 1 As shown, the specific solutions of the embodiments of the present invention are as follows:

[0056] Example 1: A method for predicting the toxicity of commercial pesticides to earthworms based on machine learning, comprising the following steps:

[0057] S1. Obtain a dataset; the dataset includes: toxicity data of commercial pesticides to earthworms, molecular characterization characteristics of commercial pesticide technicals, types of commercial pesticide adjuvants, and soil physical and chemical properties;

[0058] It should be understood that the commercial pesticides mentioned above are composed of technical pesticides and adjuvants. The molecular characteristics of commercial pesticide technicals refer to the descriptors used to characterize the molecular characteristics of the pesticide technicals. The dataset includes data on the toxicity of commercial pesticides to earthworms, the molecular characteristics of a commercial pesticide technical, the types of commercial pesticide adjuvants, and the physical and chemical properties of soil.

[0059] The toxicity data of commercial pesticides to earthworms is the concentration of commercial pesticides in the soil that causes 50% of earthworms to die within 14 days. 50 value;

[0060] Commercial pesticides include: herbicides, insecticides, fungicides, acaricides, nematicides, and plant growth regulators; specifically, herbicides include: glyphosate, glyphosate ammonium, atrazine, bentazone, fluazifop-butyl, clethosin, clethodim, etc.; insecticides include: acetamiprid, imidacloprid, nitenpyram, clothianidin, beta-cypermethrin, bifenthrin, cypermethrin, cypermethrin, etc.; fungicides include: propiconazole, difenoconazole, epoxiconazole, thiazolinone, dimethomorph, azoxystrobin, thiabendazole, etc.; acaricides include: clofentazine, hexythiazolin, fenthiophanate, tebufenpyrad, fenpyrad; plant growth regulators include: fenobucarb, diethylaluminum trichloroacetic acid; nematicides include: matrine;

[0061] Types of adjuvants for commercial pesticides include dispersible water granules, emulsions in water, suspensions, water-dispersible powders, aqueous solutions, smoke sprays, microencapsulated powders, microencapsulated suspensions, oil dispersants, effervescent tablets, seed dressings, water-soluble granules, soluble concentrates, soluble powders, concentrates, granules, wettable powders, microemulsions, and emulsifiable concentrates;

[0062] Soil physical and chemical properties include soil organic matter and soil pH. The dataset is obtained through experiments or data query. The toxicity data of commercial pesticides on earthworms in the dataset include the LC values ​​of commercial pesticides on earthworms. 50 Non-toxic data with a value of ≥100 mg / kg dry soil, LC values ​​of commercial pesticides for earthworms 50 Toxicity data with values ​​< 100 mg / kg dry soil;

[0063] The aforementioned datasets were obtained through experiments or data collection. Specifically, the experimental data (608 entries) were generated in our laboratory and the data (339 entries) were extracted from existing literature. The data were derived from a 14-day commercial pesticide toxicity test on Eisenia fetida (Eisenia fetida) in six soil types (GB / T31270.15-2014; OECD, 1984). These included five natural soils (Northeast Black Soil, Taihu Paddy Soil, Wuxi Paddy Soil, Henan Erhe Soil, and Jiangxi Red Soil) and one artificial soil prepared according to the guidelines of the Organization for Economic Cooperation and Development (OECD, 1984) with a pH of 6.0 and a SOM content of 8.5%. The remaining 339 data were obtained from public databases such as Web of Science and China National Knowledge Infrastructure (CNKI).

[0064] Data preprocessing: To obtain high-quality data, mixtures, inorganic compounds, organometallic compounds, salts, characters representing stereoisomers, duplicate data, and conflicting data were removed. The simplified molecular linear input specification (SMILES) of each commercial pesticide was collected from PubChem for molecular characterization feature calculation. Commercial pesticides without SMILES were also removed. For completely identical records (i.e., experimental conditions, SMILES, and LC 50 For records with the same values, only one was retained and the rest of the duplicates were deleted. For conflicting data, i.e., the LC values ​​corresponding to the same commercial pesticide (SMILES) under the same experimental conditions, 50 If the values ​​were different, data with consistent ranges were retained (e.g., >100 mg / kg dry soil and >200 mg / kg dry soil, >200 mg / kg dry soil was selected), and data with inconsistent ranges were removed (e.g., >100 mg / kg dry soil and >50 mg / kg dry soil). If a commercial pesticide had both ambiguous data and definite toxicity values ​​under the same experimental conditions, for commercial pesticides with more than 5 records, data outliers outside the median ±50% were removed, and the remaining LC values ​​were calculated. 50 For commercial pesticides with less than 5 records, all LC 50 The arithmetic mean of

[0065] The above-mentioned molecular characterization features include molecular descriptors and molecular fingerprints; the molecular descriptors include Mordred descriptors and RDKit descriptors, and the molecular fingerprints include Morgan fingerprints, which are calculated through SMILES encoding of pesticide technical molecules; Mordred is a Python library that supports rich MD, has fast calculation speed, and includes automatic testing functions. RDKit is an open source chemical informatics toolkit that can efficiently calculate more than 200 descriptors through its Python interface. Morgan fingerprints are one of the most commonly used fingerprints. Since binary Morgan fingerprints are only encoded as 0 and 1, the number of atomic groups cannot be quantified. Therefore, in this study, we used count-based Morgan fingerprints, which calculates the number of atomic groups at each "bit" position. The Python interface of RDKit can be used to calculate count-based Morgan fingerprints. For the calculated Mordred and RDKit descriptors, any descriptor columns containing empty values ​​or all the same values ​​are removed; for count-based Morgan fingerprints, all empty values ​​are filled with zero;

[0066] S2. Constructing a toxicity prediction model based on the data set; the input of the toxicity prediction model is the molecular characterization characteristics of the commercial pesticide technical, the type of commercial pesticide adjuvant, and the physical and chemical properties of the soil; the output of the toxicity prediction model is whether the commercial pesticide is toxic to earthworms;

[0067] Exemplarily, in a specific implementation process, the toxicity prediction model uses the QSAR model as a framework and utilizes a variety of machine learning algorithms as tools to perform model operations; in QSAR modeling, MD and MF are widely considered to be the most commonly used molecular characterization methods. MD quantitatively characterizes the physical and chemical information of chemical molecules in numerical form. MF is used to encode chemical structures, mostly expressed as a series of binary bits, indicating the presence or absence of specific substructures within the molecule. Taking into account that different types of MD and MF may affect the performance of the prediction model; the embodiment of the present invention selects two groups of MD (Mordred descriptors, RDKit descriptors) and one MF (Morgan fingerprint based on counting) for ML-QSAR modeling (QSAR model based on multiple machine learning algorithms);

[0068] First, the dataset is divided into a training set and a test set in a ratio of 8:2 through stratified sampling (it should be understood that the test set is used to evaluate the generalization ability of the model. The application method is the same as the existing technology and will not be repeated here. The following training is all completed through the training set) to ensure that the category distribution in the training set and the test set is the same as that in the original dataset.

[0069] Based on the training set, the model is trained. The model training process specifically includes the following steps S2-1 to S2-4:

[0070] S2-1. The toxic data in the data set are marked as the toxic group and the non-toxic data are marked as the non-toxic group; the toxic group is the LC value of commercial pesticides on earthworms. 50 The non-toxic data of the value ≥ 100 mg / kg dry soil is the LC value of commercial pesticides to earthworms. 50 Toxicity data with values ​​< 100 mg / kg dry soil;

[0071] For example, the dataset includes 344 non-toxic samples and 229 toxic samples, covering 200 commercial pesticides, 20 formulation types, and 17 soil types. The dataset is divided into a training set containing 458 samples and a test set containing 115 samples;

[0072] S2-2. Using an extreme gradient boosting model, key molecular descriptors or molecular fingerprints are screened from the molecular characterization features of commercial pesticide technicals. Correlation analysis is then performed to eliminate key molecular descriptors or molecular fingerprints with high autocorrelations from the key molecular descriptors or molecular fingerprints, thereby obtaining molecular characterization features after screening under three conditions, as well as corresponding commercial pesticide adjuvant types, soil physical and chemical properties, and commercial pesticide toxicity data to earthworms.

[0073] Specifically, we first used the extreme gradient boosting model to screen out three groups of molecular characterization features, retaining 10 Mordred descriptors, 10 RDKit descriptors, and 10 Morgan fingerprints for model development. Correlation analysis was used to evaluate the correlation between each group of molecular characterization features. The correlation analysis was specifically judged using the Spearman rank correlation coefficient. When the Spearman rank correlation coefficient between the molecular characterization features exceeded 0.8 and the p-value was less than 0.05, one of the molecular characterization features was randomly retained. After screening the molecular characterization feature conditions, this study selected 10 Mordred descriptors, 9 RDKit descriptors, and 9 count-based Morgan fingerprints for model training. A brief description of the 9 RDKit descriptors is shown in Table 1.

[0074] Table 1 Brief description of RDKit descriptor

[0075]

[0076] S2-3. Model training is performed using M machine learning algorithms on the molecular characterization features after the three conditions are screened, as well as the corresponding commercial pesticide adjuvant types, soil physical and chemical properties, and commercial pesticide toxicity data to earthworms, to construct M×3 candidate prediction models;

[0077] Exemplarily, classification prediction models were constructed using ten machine learning frameworks, including: LR (Logistic Regression), LDA (Linear Discriminant Analysis), KNN (K-Nearest Neighbors), GNB (Gaussian Naive Bayes), DT (Decision Tree), SVM (Support Vector Machine), Bagging (Bootstrap Aggregating), RF (Random Forest), GBDT (Gradient Boosting Decision Tree), and LightGBM (Light Gradient Boosting Machine); thus, a total of 30 classifier models were constructed by combining three types of molecular characterization with ten different machine learning algorithms.

[0078] The model inputs for each of the 30 classifier models are shown in Table 2 ;

[0079] Table 2 Model input parameters

[0080]

[0081]

[0082] S2-4. Based on the toxic group and the non-toxic group, internal validation is performed using ten-fold cross validation, and then each of the candidate prediction models is tuned using Bayesian optimization; the accuracy, recall rate, and F1 score index of the candidate prediction model are calculated, and the optimal model is screened as the toxicity prediction model based on the accuracy, recall rate, and F1 score index; specifically, the ten-fold cross validation and Bayesian optimization are conventional applications of the prior art and are not further described here;

[0083] Methods for selecting the optimal model as the toxicity prediction model based on the accuracy, recall, and F1 score indicators include:

[0084] 1) First, according to the accuracy ranking, at least 6 candidate prediction models with higher accuracy are selected from M×3 candidate prediction models;

[0085] The calculation formula of the prediction accuracy is as follows:

[0086]

[0087] In formula (2), TP represents the number of data that are actually toxic and are also predicted to be toxic by the toxicity prediction model, TN represents the number of data that are actually non-toxic and are also predicted to be non-toxic by the toxicity prediction model, FP represents the number of data that are actually non-toxic but are predicted to be toxic by the toxicity prediction model, and FN represents the number of data that are actually toxic but are predicted to be non-toxic by the toxicity prediction model.

[0088] 2) Selecting a candidate prediction model with the largest sum of recall rate and F1 score from the at least 6 candidate prediction models with high accuracy as the toxicity prediction model.

[0089] Specifically, as shown in Table 3, the confusion matrix is ​​used to illustrate the relationship between the predicted results and the actual results on the test set.

[0090] Table 3 Relationship between predicted results and actual results

[0091]

[0092]

[0093] The present invention uses accuracy, precision, recall, F1 score, specificity, Matthews correlation coefficient (MCC), and area under the receiver operating characteristic curve (AUROC) to evaluate prediction performance. Except for MCC, the value range of all indicators is 0 to 1, and the closer the value is to 1, the better the model performance. The range of MCC value is -1 to 1, with a value of 1 indicating excellent model performance and a value of -1 indicating poor model performance.

[0094] The model's precision measures the proportion of samples that the classifier predicts to be toxic and are actually toxic. It is calculated as follows:

[0095]

[0096] Recall, also known as sensitivity, represents the proportion of samples that are correctly predicted to be toxic among all samples that are actually toxic. It is calculated as follows. In this study, recall is the most important metric because accurately identifying truly toxic commercial pesticides is crucial.

[0097]

[0098] The F1 score can balance the precision and recall rate and more comprehensively evaluate the performance of the classifier. The calculation formula is as follows:

[0099]

[0100] Specificity refers to the proportion of samples that are correctly identified among those that are actually non-toxic. The calculation formula is as follows:

[0101]

[0102] The Matthews correlation coefficient (MCC) is a reliable indicator compared to other indicators. It takes into account all four categories in the confusion matrix and is calculated as follows:

[0103]

[0104] Exemplarily, the above data preprocessing, visualization, and generation of ML-QSAR models can all be completed using Python (version 3.12.2) and R (version 4.4.1) (other methods may also be used, which are not protected by the present invention);

[0105] like Figure 2 、 Figure 3 As shown in the results, all ML-QSAR models achieved high accuracy in the ten-fold cross-validation and training set. The models constructed using the Mordred descriptor, the RDKit descriptor, and the count-based Morgan fingerprint had average accuracy rates of 0.77, 0.79, and 0.76, respectively, in the test set. The metrics of the six models with relatively high accuracy are shown in Table 4. In Table 4, the GBDT model based on the RDKit descriptor performed best in terms of recall and F1 score.

[0106] Table 4 Performance of the 6 best models

[0107]

[0108]

[0109] In addition, the confusion matrices of the six models with relatively high accuracy are as follows Figure 4 As shown, Figure 4 The confusion matrix quantifies the differences in the classification capabilities of the six models, reflecting the overall accuracy and prediction details of each category;

[0110] like Figure 5 As shown in the figure, the ROC curves of six models with relatively high accuracy are constructed by plotting sensitivity (true positive rate) on the y-axis and 1-specificity (false positive rate) on the x-axis. Therefore, it is suitable for evaluating the performance of classifiers in different categories; the ROC curve can evaluate the performance of the model at different thresholds. The larger the area under the ROC curve (AUC), the better the performance. The ROC curve diagrams of 6 different models can intuitively compare the performance of these models in classification tasks, especially their ability to distinguish different categories. By comparing ROC curves, it is possible to determine which models are more effective in specific application scenarios, thereby providing a basis for selecting the best model. Figure 5 It can be seen that the RF model is more preferred.

[0111] It's important to note that accurately identifying commercial pesticides toxic to Eisenia fetida is more critical than correctly classifying non-toxic samples. Misclassifying toxic commercial pesticides as non-toxic can lead to significant environmental risks, while false positives often require only additional, but controlled, toxicity testing. Therefore, a model's ability to identify toxic commercial pesticides is central to evaluating its overall effectiveness. Consequently, the GBDT model classifier built using RDKit descriptors performed optimally, successfully predicting the toxicity of commercial pesticides to earthworms.

[0112] Before using the toxicity prediction model for prediction in actual practice, the reliability of the prediction should be determined first. Methods for determining the reliability of the prediction include:

[0113] First, determine the applicable domain threshold of the molecular characterization features of commercial pesticide technicals in the toxicity prediction model;

[0114] Calculate the average Euclidean distance between the molecular characteristics of the commercial pesticide technical in the input of the toxicity prediction model and the molecular characteristics of the commercial pesticide technical in the nearest neighbor sample in the training set;

[0115] When the average Euclidean distance is less than or equal to the applicable domain threshold, the prediction result is considered reliable; conversely, when the average Euclidean distance is greater than the applicable domain threshold, the prediction result is considered unreliable.

[0116] The calculation formula of the applicable domain threshold is as follows:

[0117]

[0118] In formula (1), is the average Euclidean distance between the molecular characterization feature of each commercial pesticide technical in the training set and its three nearest neighbors in the dataset, σ is the standard deviation of the Euclidean distances between the three nearest neighbors, and Z is a constant ranging from 0.4 to 0.6 (0.5 is selected in this example).

[0119] For example, the result of judging the reliability of the prediction is as follows: Figure 6 As shown, the applicability domain of the GBDT model built using the RDKit descriptor is defined by the k-nearest neighbor method with a threshold set to 35 ( Figure 6 (a)); all samples in the test set are included in the applicability domain ( Figure 6 (b)), which indicates that the model's prediction results are reliable; Figure 6 In the SHAP bee swarm graph, the x-axis represents the index of each sample in the training set. The y-axis represents the average Euclidean distance between each sample in the training set and its three nearest neighbors. The color of the point represents the size of the Euclidean distance. The features in the SHAP bee swarm graph are sorted from top to bottom according to the average absolute SHAP value of each feature ( Figure 6(c)); Each point in the figure represents a sample. For samples with the same value, slight perturbations are added in the vertical direction to avoid overlap and enhance readability; the x-axis represents the calculated SHAP value, where the positive axis indicates the prediction as the positive class (i.e., toxic commercial pesticides) and the negative axis indicates the prediction as the negative class (i.e., non-toxic commercial pesticides);

[0120] In an embodiment of the present invention, a method for analyzing the most important input characteristic factors affecting the toxicity prediction model based on SHAP analysis is also provided. The input characteristic factors obtained from the analysis are used to adjust or reduce these inputs in subsequent practical processes, resulting in commercial pesticides with lower toxicity to earthworms. The results show that the properties of commercial pesticides represented by the RDKit descriptor can significantly affect the toxicity of chemicals to Eisenia fetida. Combined with the content shown in Table 1 above, FpDensityMorgan2 represents the Morgan fingerprint density within the two-bond range around each atom and has the greatest impact on the model. Commercial pesticides with higher FpDensityMorgan2 values ​​mean that they have more unique substructures or functional groups within a specified radius and are more likely to be toxic. PEOE_VSA6 represents partial charge averaging, which calculates the partial charges of each atom in the molecule. Higher charges indicate higher toxicity. BCUT2D_CHGHI represents the charge (CHG) high value (HI) value. Higher values ​​indicate lower toxicity. Commercial pesticide technical is a mixture of active ingredients and related impurities produced during the manufacturing process. Because technical contains no or only a small amount of essential additives, it has lower toxicity to Eisenia fetida. Wettable powders consist of technical powder, fillers, dispersants, and surfactants, and do not use solvents or emulsifiers, making them safer for the environment. Microemulsions consist of oil, water, and surfactants, sometimes with the addition of co-surfactants, and contain little or no organic reagents, so they put less pressure on the environment. Unlike microemulsions, emulsifiable concentrates and aqueous emulsions are liquids formed by adding emulsifiers. They are usually mixtures of nonionic or anionic surfactants with solvents (such as xylene, methylnaphthalene, or other petroleum solvents) and alcohol solvents (such as methyl isobutyl ketone or isopropyl alcohol), resulting in higher toxicity to Eisenia fetida. At the same time, although the organic solvent content in water-dispersible granules and suspensions is reduced, the complexity of the adjuvant formulation increases, which may lead to higher toxicity to Eisenia fetida.

[0121] Soil organic matter content is a significant factor influencing the toxicity of chemicals in the soil environment to Eisenia fetida. In soils with high soil organic matter content, commercial pesticides may be more easily degraded, exhibit higher adsorption affinity, and have lower mobility, ultimately reducing their toxicity to Eisenia fetida. Furthermore, pH also affects the toxicity of chemicals in the soil environment. At acidic pH, the degradation rate of chemicals is slower than at alkaline and neutral pH, and the stability of various chemical groups increases at acidic pH, which can increase the toxicity of commercial pesticides to Eisenia fetida.

[0122] S3. Based on the toxicity prediction model, whether commercial pesticides are toxic to earthworms is predicted. In order to optimize the model performance, Bayesian optimization is used to tune the model hyperparameters. The results show that the toxicity prediction model (i.e., ML-QSAR model) shows high prediction accuracy in both the training set and the test set.

[0123] In summary, the toxicity of mixed chemical substances is affected by multiple factors. In this example, commercial pesticides (original pesticides and adjuvants) are used to represent mixed chemical substances, and soil non-target organisms, earthworms, are used to develop a method and model for predicting the biological toxicity of commercial pesticides in a real environment. The results show that the GBDT model classifier constructed using the RDKit descriptor performs best and can successfully predict the toxicity of commercial pesticides to earthworms. In the test set, it achieved an accuracy of 0.83, a precision of 0.77, a recall rate of 0.80, an F1 score of 0.79, an MCC of 0.64, and an AUROC of 0.88. SHAP analysis revealed the morphology and molecular structure characteristics (Morgan fingerprint density, molecular charge distribution) of mixed chemical substances. Soil environmental factors, soil organic matter content, and pH are important factors affecting the toxicity of mixed chemical substances. The results show that the toxicity of commercial pesticides in a real environment can be successfully predicted using machine learning models, which enhances the ability to explain the toxicity mechanism of chemical substances in a real environment and provides strong support for supplementing and improving the basic data of chemical substances.

Claims

1. A method for predicting the toxicity of commercial pesticides to earthworms based on machine learning, characterized in that: The following steps are involved: Acquire a data set; the data set includes: toxicity data of commercial pesticides to earthworms, molecular characterization characteristics of commercial pesticide technicals, types of commercial pesticide adjuvants, and soil physical and chemical properties; Constructing a toxicity prediction model based on the data set; the input of the toxicity prediction model is the molecular characterization characteristics of the commercial pesticide technical, the type of commercial pesticide adjuvant, and the physical and chemical properties of the soil; the output of the toxicity prediction model is whether the commercial pesticide is toxic to earthworms; Based on the toxicity prediction model, predict whether commercial pesticides are toxic to earthworms.

2. The method according to claim 1, wherein The toxicity data of the commercial pesticides to earthworms is the concentration of the commercial pesticide in the unit soil that causes 50% of the earthworms to die within 14 days. 50 value.

3. The method according to claim 1, wherein The commercial pesticides include: herbicides, insecticides, fungicides, acaricides, nematicides, and plant growth regulators; Types of adjuvants for commercial pesticides include dispersible water granules, emulsions in water, suspensions, water-dispersible powders, aqueous solutions, smoke sprays, microencapsulated powders, microencapsulated suspensions, oil dispersants, effervescent tablets, seed dressings, water-soluble granules, soluble concentrates, soluble powders, concentrates, granules, wettable powders, microemulsions, and emulsifiable concentrates.

4. The method according to claim 1, wherein The soil properties include soil organic matter and soil pH.

5. The method according to claim 1, wherein The dataset is obtained through experiments or data query. The toxicity data of commercial pesticides on earthworms in the dataset include LC 50 Non-toxic data with a value of ≥100 mg / kg dry soil, LC values ​​of commercial pesticides for earthworms 50 Toxicity data with values ​​< 100 mg / kg dry soil.

6. The method according to claim 1, wherein The molecular characterization features include molecular descriptors and molecular fingerprints; the molecular descriptors include Mordred descriptors and RDKit descriptors, and the molecular fingerprints include Morgan fingerprints.

7. The method according to claim 6, wherein The method for constructing a toxicity prediction model based on the data set includes: Mark the toxic data in the data set as the toxic group and the non-toxic data as the non-toxic group; The extreme gradient boosting model was used to screen key molecular features from the molecular features of commercial pesticide technicals. Correlation analysis was then performed to eliminate key molecular features with high autocorrelation, resulting in the screening of molecular features and the corresponding commercial pesticide adjuvant types, soil physical and chemical properties, and toxicity data of the commercial pesticides to earthworms. Model training is performed using M machine learning algorithms on the molecular characterization features after the condition screening, the corresponding commercial pesticide adjuvant types, soil physical and chemical properties, and commercial pesticide toxicity data to earthworms, to construct M×3 candidate prediction models; Based on the toxic group and the non-toxic group, ten-fold cross validation was used for internal validation, and then Bayesian optimization was used to tune each of the candidate prediction models; the accuracy, recall rate and F1 score indicators of the candidate prediction models were calculated, and the optimal model was selected as the toxicity prediction model based on the accuracy, recall rate and F1 score indicators.

8. The method according to claim 7, wherein The method for selecting the optimal model as the toxicity prediction model based on the accuracy, recall rate, and F1 score indicators includes: First, according to the accuracy ranking, at least 6 candidate prediction models with higher accuracy are selected from M×3 candidate prediction models; Among the at least six candidate prediction models with relatively high accuracy, a candidate prediction model with the largest sum of recall rate and F1 score is selected as the toxicity prediction model.

9. The method according to claim 7, wherein: Before making a prediction based on a toxicity prediction model, the reliability of the prediction should be determined first. Methods for determining the reliability of the prediction include: First, determine the applicable domain threshold of the molecular characterization features of commercial pesticide technicals in the toxicity prediction model; Calculate the average Euclidean distance between the molecular characteristics of the commercial pesticide technical in the input of the toxicity prediction model and the molecular characteristics of the commercial pesticide technical in the nearest neighbor sample in the training set; When the average Euclidean distance is less than or equal to the applicable domain threshold, the prediction result is considered reliable; conversely, when the average Euclidean distance is greater than the applicable domain threshold, the prediction result is considered unreliable.

10. The method according to claim 9, wherein The calculation formula of the applicable domain threshold is as follows: In formula (1), is the average Euclidean distance between the molecular representation features of each commercial pesticide technical in the training set and its three nearest neighbors, σ is the standard deviation of the Euclidean distances between the three nearest neighbors, and Z is a constant ranging from 0.4 to 0.6.