Biotoxicity prediction method based on SMOTE small sample machine learning

The virtual data is generated through SMOTE small sample machine learning method, which solves the problem of scarcity of nanotoxicology data in natural water bodies, and realizes accurate prediction of the accumulation of nanometals in aquatic organisms under limited samples, expands the scope of toxicity prediction, and provides a new method for rapid risk assessment and monitoring.

CN120452585APending Publication Date: 2025-08-08BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510385238.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

In the nanotoxicology research of natural water bodies, existing machine learning methods are prone to overfitting or underfitting due to complex experimental conditions and scarce data, and the prediction results are inaccurate, making it difficult to effectively use limited data for biotoxicity prediction.

Method used

SMOTE small sample machine learning method is adopted to generate virtual data by synthesizing a few types of oversampling technology, supplement sample data and maintain sample balance, and combine data preprocessing and feature importance analysis to build a robust machine learning model, including data collection, preprocessing, clustering analysis, model verification and feature importance analysis.

Benefits of technology

Accurate prediction of the accumulation of nanometals in aquatic organisms under limited samples, expand the scope of toxicity prediction, can process complex data, save experimental costs, provide new methods for rapid risk assessment and monitoring, and solve the limitations of small sample data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452585A_ABST
    Figure CN120452585A_ABST
Patent Text Reader

Abstract

A biotoxicity prediction method based on SMOTE small sample machine learning can deal with the problem of insufficient natural experimental data, is proved to be capable of solving complex environmental problems, and comprises the steps of generating virtual data based on a synthetic minority class oversampling technology SMOTE to supplement sample data and maintain sample balance to meet the machine learning requirement; the method comprises the following steps: collecting natural water experimental data to obtain data related to metal toxicity prediction: water chemical condition data, nano-metal data and biological accumulation data; the data preprocessing is used for converting the data into a form suitable for being input into a machine learning model so as to improve the precision and efficiency of a fitting model; constructing a small sample machine learning model generated based on the virtual sample; model verification and feature importance analysis are carried out, the model verification comprises internal and external verification on the model so as to correctly evaluate the robustness and predictive ability of the model, and the feature importance analysis comprises feature importance sorting by adopting two standards of mean square error (MSE) increase and node purity increase.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of toxicity prediction of pollutants in water environments to aquatic organisms, and in particular to a biological toxicity prediction method based on SMOTE small sample machine learning. Background Art

[0002] With the rapid development of global industrialization and modernization, a large number of new chemicals are being produced and released, and biotoxicity has become a widespread concern. Biotoxicity refers to the toxic effects of chemicals on organisms, particularly aquatic and terrestrial organisms. These chemicals may originate from industrial wastewater, agricultural residues, pharmaceuticals, and personal care products. They enter ecosystems through various pathways, posing a serious threat to biodiversity and ecological balance. Therefore, accurately predicting and assessing the biotoxicity of these chemicals is of great significance for environmental protection and ecological security.

[0003] Currently, traditional methods for toxicity prediction mostly use experimental research, traditional linear regression models, biological ligand models, etc. The experiments are complex, long-term and costly, making it difficult to accurately reveal the toxicity mechanism and environmental behavior of nanomaterials as quickly as possible. In comparison, the current research on machine learning methods has a rapid processing process and more accurate prediction results. Therefore, machine learning methods are considered to be an emerging and effective quantitative means to accurately assess and predict the environmental risks of nanoparticles, and can be used to predict the environmental behavior and toxicity of nanoparticles.

[0004] However, current research using machine learning methods is based on the premise that data is sufficient. However, due to the diverse experimental conditions and numerous complex operations in nanotoxicology conducted in natural water bodies, there is less homogeneous data available for building machine learning models. If traditional machine learning modeling methods are still used, the model will fall into the error of overfitting or underfitting, increasing the model error, widening the gap between the predicted value and the true value, increasing the inaccuracy of the results, and even failing to meet the model requirements, making it impossible to proceed to the next step of research. Scarcity remains a major bottleneck in biological toxicity prediction. How to effectively utilize limited data and improve the performance of prediction models has become a hot topic and difficulty in current research.

[0005] List of comparison documents for existing solutions:

[0006] 1) Chemical Status: Zhu, L., et al. (2018). QSAR modeling of chronic toxicity of organic pollutants to aquatic organisms. Chemosphere, 212, 347-354.

[0007] Panagiotis Isigonis, Antreas Afantitis, Dalila Antunes et al. RiskGovernance ofEmerging Technologies Demonstrated in Terms of its Applicability to Nanomaterials[J].Small.2020.16(36):2003303.

[0008] Yan Yupeng, Tang Yadong, Wan Biao, et al. Effect of particle size on environmental behavior of nano-oxides[J]. Environmental Science. 2018. 39(06): 2982-2990.

[0009] 2) Application of machine learning models that require large amounts of data: D.A. Winkler. Role of Artificial Intelligence and Machine Learning in Nanosafety[J]. Small. 2020.(36):e2001883.

[0010] S.Cipullo,B.Snapir,G.Prpichet al.Prediction of bioavailability and toxicity of complex chemical mixtures through machine learning models[J].Chemosphere.2019.215:388-395.

[0011] Y.Zhou,Y.Wang,W.Peijnenburget al.Using Machine Learning to PredictAdverse Effects of Metallic Nanomaterials to Various Aquatic Organisms[J].Environ Sci Technol.2023. Summary of the Invention

[0012] This paper addresses the shortcomings of existing technologies by providing a method for predicting biological toxicity based on small-sample machine learning using SMOTE. This method can process complex data from various dimensions, effectively reducing experimental costs and enabling rapid predictions, thus providing a new approach for risk assessment and monitoring. This method comprehensively considers the impact of pollutants, environmental conditions, and species on toxicity, effectively addressing the issue of small sample sizes and overcoming the limitations of sample data.

[0013] The technical solutions of the present invention are as follows:

[0014] A method for predicting biological toxicity based on SMOTE small sample machine learning is characterized by generating virtual data based on the synthetic minority oversampling technique SMOTE to supplement sample data and maintain sample balance to meet machine learning requirements. The steps are as follows:

[0015] Step 1: Collect and screen natural water experimental data to obtain the following data related to predicting the accumulation of nanometals in organisms: water chemical condition data, nanometal data, and bioaccumulation data;

[0016] Step 2: Data preprocessing to convert the data into a form suitable for input into the machine learning model to improve the accuracy and efficiency of the fitting model;

[0017] Step 3: cluster analysis of real data to avoid blindly generating virtual data and resulting in an unbalanced dataset;

[0018] Step 4: Build a SMOTE small sample machine learning model;

[0019] Step 5: Model validation and feature importance analysis. The model validation includes internal and external validation of the SMOTE small sample machine learning model to correctly evaluate its robustness and predictive ability, thereby selecting the optimal model. The feature importance analysis includes ranking the feature importance using two criteria: increasing mean square error (MSE) and increasing node purity.

[0020] Step 1 includes the water quality obtained in the experiment, the size of the nano-gold particles synthesized by ourselves, and the toxic effects of nano-gold in natural water in the body; Step 1 includes the following data classification: the bioaccumulation data is used as the predictor variable, and the rest of the data are characteristic variables; the water chemical condition data in step 1 include water temperature, hardness, pH, total organic carbon TOC content and total inorganic carbon content, anion F in water - 、Cl - 、NO2 - 、NO3 - 、SO4 2- and CO3 2- The content of metal cations Na in water + , K + , Ca 2+ and Mg 2+ The content of trace metals Cd, Zn, Cu, Ni, Ba, Cr, Mn, As, Pb and Se in water at the ppm level.

[0021] Step 2 includes zero-variance variable detection to remove noisy data or features, feature selection based on correlation analysis to remove redundant variables, and z-score normalization to improve the comparability of data between different features.

[0022] Step 3 includes selecting the minimization of the within-cluster sum of squares method in K-means clustering to perform cluster analysis on the real data before generating virtual data to understand the distribution of the real data in different categories.

[0023] It includes using SMOTE to generate new samples, namely virtual samples, between minority class samples through random interpolation method. For each minority class sample, the distance between it as the current sample and all other minority class samples is calculated, and K nearest neighbor distances are found, where K is a positive integer. A nearest neighbor sample is randomly selected from these K nearest neighbor distances, and the difference between the nearest neighbor sample and the current sample is calculated. Then, based on the difference ratio, a new synthetic sample is generated. The synthetic sample, namely the virtual sample, is located on the line connecting the nearest neighbor sample and the current sample. The generated virtual samples are screened using the following relative error equation, and virtual samples with a relative error of less than 10% are confirmed as valid virtual samples:

[0024]

[0025] where Y i Represents the output variable, Y i ′ represents the target output variable. In practical applications, i ′, firstly, an initial model is constructed using real experimental data, and then the virtual input data is substituted into the initial model to obtain an approximate target output data as the real Y i ′, ||·|| represents the modulus of the vector;

[0026] After the generated virtual samples are screened, they are aggregated with the original data to form the complete data set used to build the SMOTE small sample machine learning model.

[0027] The complete data set is divided into a training set with a data volume of 80% and a test set with a data volume of 20% through stratified sampling, and the test set contains more real data than virtual data.

[0028] Step 4 includes using the following five machine learning models to build a relationship model between influencing factors and bioaccumulation: classification regression tree CART, random forest RF, gradient boosting machine GBM, support vector machine SVM, artificial neural network ANN; using random search method for all relationship models to cross-validate the correlation coefficient Q 2 CV The optimal hyperparameters of the model are determined by using the

[0029] The step 5 includes the following formula:

[0030]

[0031] where Q 2 CV represents the correlation coefficient of cross validation, Q 2 ext represents the external validation coefficient of determination, the subscripts CV and ext represent cross validation and external validation, respectively, RMSE CV或ext Represents the cross-validation or external validation root mean square error, MAE CV represents the cross-validation mean absolute error, R 2 Represents the correlation coefficient of the training set, y j Indicates the acute toxicity observation value, Indicates the acute toxicity prediction value, represents the average value of all acute toxicity observations, i represents the sample number, and n represents the total number of samples.

[0032] The technical effects of the present invention are as follows: The present invention provides a biological toxicity prediction method based on SMOTE small sample machine learning, which can be used to study the accumulation of nanometals in aquatic organisms in natural water bodies. This method comprehensively considers the effects of natural water environment media and nanometal environmental behavior on toxicity, solves the limitation of machine learning being limited by experimental data, and can realize the prediction of metal toxicity to aquatic organisms under different water chemical conditions under limited samples. Then, feature importance analysis is performed based on the support vector machine model to obtain the key factors affecting toxicity.

[0033] Features and advantages of the present invention: 1) Comprehensive consideration of natural water chemical conditions, the physicochemical properties of metals, the environmental behavior of nanometals in natural water bodies, and the accumulation of nanometals in organisms. Compared with existing toxicity prediction methods, the scope of toxicity prediction has been expanded, and the toxicity of unknown metals can be predicted, as well as the toxicity of unknown species. 2) The small sample machine learning method generated based on virtual samples can process complex data of different dimensions, effectively save experimental costs, and make predictions quickly, providing a new method for risk assessment and monitoring. Compared with traditional statistical methods, it effectively solves the problem of small samples and breaks the limitations of sample data. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 The present invention is a flow chart of a method for predicting biological toxicity based on small sample size machine learning using SMOTE (Synthetic Minority Over-sampling Technique). Figure 1The method includes step 1, natural water experimental data collection to obtain the following data related to the prediction of metal toxicity data: water chemical condition data, nanometal data, and bioaccumulation data; step 2, data preprocessing to convert the data into a form suitable for input into the machine learning model to improve the accuracy and efficiency of the fitting model; step 3, building a small sample machine learning model; step 4, model validation and feature importance analysis, the model validation includes internal and external validation of the model to correctly evaluate its robustness and predictive ability, so as to select the optimal model, and the feature importance analysis includes using two criteria: increasing mean square error (MSE) and increasing node purity to rank the feature importance.

[0035] Figure 2 The cross-validation results of the bioaccumulation prediction models using virtual data at each multiple of the optimal model. (A) Results for the RF model using 10-fold virtual data. (B) Results for the SVM model using 12-fold virtual data. (C) Results for the SVM model using 14-fold virtual data. (D) Results for the SVM model using 16-fold virtual data. (E) Results for the SVM model using 18-fold virtual data. (F) Results for the RF model using 20-fold virtual data. The 14-fold model was the optimal. Figure 2 LM is the Lagrange Multiplier model (LM), CART is the classification and regression tree, RF is the random forest, GBM is the gradient boosting machine, SVM is the support vector machine, and ANN is the artificial neural network; MAE represents the cross-validation mean absolute error, RMSE represents the cross-validation or external validation root mean square error, and Rsquared is R 2 Represents the model results and external validation coefficient of determination. DETAILED DESCRIPTION

[0036] Below is the attached figure ( Figure 1-Figure 2 ) and Examples illustrate the present invention.

[0037] Figure 1 It is a flow chart of the method for implementing the present invention of small sample machine learning in biological toxicity prediction under data scarcity. Figure 2 This is the cross-validation result of each bioaccumulation prediction model under the virtual data of each multiple optimal model. Figures 1 to 2 As shown in the figure, a method for predicting biological toxicity using small-sample machine learning under data scarcity is characterized by generating virtual data based on the synthetic minority oversampling technique (SMOTE), supplementing sample data, and maintaining sample balance to meet machine learning requirements. This method improves prediction accuracy, reduces model constraints, and improves generalization. The method includes the following steps:

[0038] Step 1: Collect and screen natural water experimental data to obtain the following data relevant to predicting the accumulation of nanometals in organisms: water chemistry data, nanometal data, and bioaccumulation data; experimental water quality, the size of self-synthesized nanogold particles, and the toxic effects of nanogold in natural water in vivo;

[0039] Step 2: Data preprocessing to convert the data into a form suitable for input into the machine learning model to improve the accuracy and efficiency of the fitting model;

[0040] Step 3: Cluster analysis of real data to avoid blindly generating virtual data and resulting in an unbalanced dataset;

[0041] Step 4: Build a SMOTE small sample machine learning model

[0042] Step 5: Model validation and feature importance analysis. The model validation includes internal and external validation of the model to correctly evaluate its robustness and predictive ability, thereby selecting the optimal model. The feature importance analysis includes ranking the feature importance using two criteria: increasing mean square error (MSE) and increasing node purity.

[0043] The water quality in step 1 includes: water temperature, hardness, pH, total organic carbon content, total inorganic carbon content and TOC, etc.; - 、Cl - 、NO2 - 、NO3 - 、SO4 2- 、CO3 2- The content of anions, Na + , K + , Ca 2+ Mg 2+ The content of metal cations with higher concentrations (ppm level) in water and the concentrations of trace metals such as Cd, Zn, Cu, Ni, Ba, Cr, Mn, As, Pb, and Se.

[0044] The step 1 includes the following data classification: the bioaccumulation data is used as the predictor variable, and the remaining data are all characteristic variables.

[0045] The data preprocessing in step 2 includes zero-variance variable detection to remove noisy data or features, feature selection based on correlation analysis to remove redundant variables, and z-score standardization to improve the comparability of data between different features (z-score is a z score or standard score used to measure the distance between a data point and the mean).

[0046] The cluster analysis in step 3 selects K-means clustering (minimizing the intra-cluster sum of squares method) to perform cluster analysis on the real data before generating virtual data, so as to understand the distribution of real data in different categories and prevent the blind generation of virtual data, which is likely to lead to an unbalanced data set.

[0047] The cluster analysis in step 3 uses K-means clustering (minimizing the within-cluster sum of squares). Before generating virtual data, cluster analysis is performed on the real data to understand the distribution of the real data across different categories. This prevents blindly generating virtual data, which is likely to result in an unbalanced dataset. K centers are randomly selected, and each data point is assigned to the cluster with the nearest center, and the cluster center is updated. This process is iterated until the cluster center no longer changes or the maximum number of iterations is reached, ultimately minimizing the sum of the squared distances from the data point to the cluster center.

[0048] It involves using SMOTE to generate new samples (SMOTE, Synthetic Minority Over-sampling Technique, synthetic minority oversampling technology) between minority class samples through random interpolation. For each minority class sample, the distance between it and all other minority class samples is calculated, and its K nearest neighbor distances are found. A sample is randomly selected from these K nearest neighbor distances, and the difference between this sample and the current sample is calculated. Then, based on the difference ratio, a new synthetic sample is generated, which is located on the line connecting the two samples. The generated virtual samples still need to be screened. When the relative error is less than 10%, the virtual sample is acceptable. Therefore, virtual samples with a relative error of less than 10% are considered valid virtual samples, and their relative error equation is:

[0049]

[0050] where Y i Represents the output variable, Y i ′ represents the target output variable. However, in practical applications, we cannot get the real Y i ′. Therefore, we first constructed an initial model using real experimental data. We then substituted virtual input data into the initial model to obtain an approximate target output data, where ∥·∥ represents the modulus of the vector. After the screening is complete, the data is combined with the original data to form the complete dataset for model construction.

[0051] In step 4, the preprocessed and virtualized data set is divided into a training set (80%) and a test set (20%) through stratified sampling. During stratified sampling, real data and virtual data are separated, and a larger proportion of real data is extracted into the test set.

[0052] Step 4 includes using the following five machine learning models to build a relationship model between feature variables and predictor variables: classification regression tree CART, random forest RF, gradient boosting machine GBM, support vector machine SVM, artificial neural network ANN; using random search method for all relationship models to cross-validate the correlation coefficient Q 2 The optimal hyperparameters of the model are determined by using the

[0053] After the model is established, it needs to be internally and externally validated to correctly evaluate its robustness and predictive ability, so as to select the optimal model. The training set is used for model training and internal validation, and the test set is used for external validation. Internal validation is achieved by three k (here k is 10) fold cross validation methods. K-fold cross validation is a technique commonly used to estimate performance and prevent overfitting. A random search method is used for all relationship models to cross-validate the correlation coefficient Q. 2 The following formulas are used for internal validation and external validation in step 5:

[0054]

[0055] where Q 2 CV represents the correlation coefficient of cross validation, Q 2 ext represents the external validation coefficient of determination, the subscripts CV and ext represent cross validation and external validation, respectively, RMSE CV或ext Represents the cross-validation or external validation root mean square error, MAE CV represents the cross-validation mean absolute error, R 2 Represents the correlation coefficient of the training set, y i Indicates the observed value of acute toxicity, represents the predicted value of acute toxicity, represents the average value of all acute toxicity observations. i represents the sample number, and n represents the total number of samples.

[0056] refer to Figure 2First, based on real experimental data, virtual data with 10-fold, 12-fold, 14-fold, 16-fold, 18-fold, and 20-fold expansion were generated. Four machine learning algorithms were used to construct prediction models for gold nanoparticle accumulation concentration. Comparisons with the traditional statistical method (LM) showed that the model fit was best when generating 16-fold virtual data, with the optimal model being the SVM. This model significantly outperformed the LM model across all evaluation criteria, demonstrating the superiority of small-sample machine learning algorithms based on virtual sample generation in processing high-dimensional nonlinear data. Then, based on real experimental data, virtual data with 10-fold, 12-fold, 14-fold, 16-fold, 18-fold, and 20-fold expansion were generated. Five machine learning algorithms were used to construct models related to water chemical conditions and gold nanoparticle accumulation. Comparisons with the traditional statistical method (LM) showed that the model fit was best when generating 14-fold virtual data, with the optimal model being the SVM. This model significantly outperformed the LM model across all evaluation criteria, demonstrating the superiority of this algorithm in processing high-dimensional nonlinear data, breaking the limitation of different sample data dimensions.

[0057] The feature importance analysis in step 4 is performed based on the SVM model. To avoid the deviation caused by using a single measurement indicator, the feature importance is ranked using the increase in mean square error (MSE) and the increase in node purity.

[0058] The method of small-sample machine learning based on SMOTE method to predict nanomaterial toxicity can address the problem of insufficient natural experimental data and has been proven to be able to solve complex environmental problems. It is characterized by including step 1, natural water experimental data collection to obtain the following data related to the prediction of metal toxicity data: water chemical condition data, nanometal data, and bioaccumulation data; step 2, data preprocessing to convert the data into a form suitable for input into the machine learning model to improve the accuracy and efficiency of the fitting model; step 3, constructing a small-sample machine learning model generated based on virtual samples; step 4, model validation and feature importance analysis, the model validation includes internal and external validation of the model to correctly evaluate its robustness and predictive ability, so as to select the optimal model, and the feature importance analysis includes using two criteria of increasing mean square error (MSE) and increasing node purity to rank feature importance.

[0059] For a long time, the great challenges faced by machine learning methods are the small number of samples, data imbalance and poor interpretability, which affect the predictive credibility of the model. Although high-precision prediction or classification tasks can be achieved by configuring appropriate parameters, the internal operation of the model is vague, and the poor interpretability also blurs the causal relationship. Therefore, appropriate feature importance analysis is very necessary. The feature importance analysis of the present invention is based on the SVM model under the most virtual multiples. In order to avoid the deviation caused by the use of a single measurement indicator, the two criteria of MSE increase and node purity increase are used to rank the feature importance. The increase in MSE is based on the decrease in forest prediction accuracy after variable perturbation; the increase in node purity is based on the change in node purity after variable splitting (measured by the residual sum of squares).

[0060] Nanomaterials constitute a major portion of the chemical industry. With the rapid development of fields such as materials and manufacturing, microelectronics and computer information technology, energy, environment, and health, countless nanomaterials have been detected in natural water environments, of which nanometals account for 80%. Their tiny size makes them more easily absorbed by organisms. Their toxicity mechanisms are complex and can be genotoxic, even affecting biomes and food chains. Therefore, the toxicity prediction and assessment of various nanomaterials is urgently needed. Therefore, this study selected nanometals as the research subject.

[0061] “Nanomaterial (particle)” refers to a natural or artificial material in the form of loose or agglomerated solid particles, which account for more than 50% of the total number of particles in the entire material and meet at least one of the following conditions: (a) one or more external dimensions of the particles are between 1nm and 100nm; (b) the particles are elongated, such as rods, fibers or tubes, with two external dimensions (diameters) less than 1nm and the other dimension (length) greater than 100nm; (c) the particles are flake-shaped, with one external dimension (thickness) less than 1nm and the other dimension greater than 100nm.

[0062] To accurately predict the toxicity of nanomaterials to aquatic species under varying water chemistry conditions, without being limited by the quantity and type of experimental data, this paper proposes a small-sample machine learning method for toxicity prediction based on virtual sample generation. This method can process complex data from diverse dimensions, effectively reducing experimental costs and enabling rapid predictions, thus providing a new approach for risk assessment and monitoring. This method comprehensively considers the impact of pollutants, environmental conditions, and species on toxicity, effectively addressing the issue of small sample sizes and breaking the limitations of sample data.

[0063] Any content not described in detail in this specification is prior art known to those skilled in the art. It should be noted that the above description is intended to help those skilled in the art understand the present invention, but does not limit the scope of protection of the present invention. Any equivalent substitution, modification, improvement, and / or simplification of the above description that does not depart from the essence of the present invention shall fall within the scope of protection of the present invention.

Claims

1. A biological toxicity prediction method based on SMOTE small sample machine learning, characterized in that: This involves generating virtual data based on the synthetic minority oversampling technique SMOTE to supplement sample data and maintain sample balance to meet machine learning requirements. The steps are as follows: Step 1: Collect and screen natural water experimental data to obtain the following data related to predicting the accumulation of nanometals in organisms: water chemical condition data, nanometal data, and bioaccumulation (biotoxicity) data; Step 2: Data preprocessing to convert the data into a form suitable for input into the machine learning model to improve the accuracy and efficiency of the fitting model; Step 3: cluster analysis of real data to avoid blindly generating virtual data and resulting in an unbalanced dataset; Step 4: Build a SMOTE small sample machine learning model; Step 5: Model validation and feature importance analysis. The model validation includes internal and external validation of the SMOTE small sample machine learning model to correctly evaluate its robustness and predictive ability, thereby selecting the optimal model. The feature importance analysis includes ranking the feature importance using two criteria: increasing mean square error (MSE) and increasing node purity.

2. The biological toxicity prediction method based on SMOTE small sample machine learning according to claim 1 is characterized in that: Step 1 includes the experimentally obtained water quality, the size of the synthesized gold nanoparticles, and the toxic effects of gold nanoparticles in natural water on the body. Step 1 includes the following data classification: the bioaccumulation data is used as the predictor variable, and the remaining data are all characteristic variables; The water chemical condition data in step 1 include water temperature, hardness, pH, total organic carbon TOC content and total inorganic carbon content, anion F in water - 、Cl - 、NO2 - 、NO3 - 、SO4 2- and CO3 2- The content of metal cations Na in water + , K + , Ca 2+ and Mg 2+ The content of trace metals Cd, Zn, Cu, Ni, Ba, Cr, Mn, As, Pb and Se in water at the ppm level.

3. The biological toxicity prediction method based on SMOTE small sample machine learning according to claim 1 is characterized in that: Step 2 includes zero-variance variable detection to remove noisy data or features, feature selection based on correlation analysis to remove redundant variables, and z-score normalization to improve the comparability of data between different features.

4. The biological toxicity prediction method based on SMOTE small sample machine learning according to claim 1, characterized in that: Step 3 includes selecting the minimization of the within-cluster sum of squares method in K-means clustering to perform cluster analysis on the real data before generating virtual data to understand the distribution of the real data in different categories.

5. The biological toxicity prediction method based on SMOTE small sample machine learning according to claim 1, characterized in that: It includes using SMOTE to generate new samples, namely virtual samples, between minority class samples through random interpolation method. For each minority class sample, the distance between it as the current sample and all other minority class samples is calculated, and K nearest neighbor distances are found, where K is a positive integer. A nearest neighbor sample is randomly selected from these K nearest neighbor distances, and the difference between the nearest neighbor sample and the current sample is calculated. Then, based on the difference ratio, a new synthetic sample is generated. The synthetic sample, namely the virtual sample, is located on the line connecting the nearest neighbor sample and the current sample. The generated virtual samples are screened using the following relative error equation, and virtual samples with a relative error of less than 10% are confirmed as valid virtual samples: where Y i Represents the output variable, Y i Indicates the target output variable. In practical applications, for Y i First, an initial model is constructed using real experimental data, and then the virtual input data is substituted into the initial model to obtain an approximate target output data as the real Y′ i , ∥·∥ represents the modulus of the vector; After the generated virtual samples are screened, they are aggregated with the original data to form the complete data set used to build the SMOTE small sample machine learning model.

6. The biological toxicity prediction method based on SMOTE small sample machine learning according to claim 5, characterized in that: The complete data set is divided into a training set with a data volume of 80% and a test set with a data volume of 20% through stratified sampling, and the test set contains more real data than virtual data.

7. The biological toxicity prediction method based on SMOTE small sample machine learning according to claim 1, characterized in that: Step 4 includes using the following five machine learning models to build a relationship model between influencing factors and bioaccumulation: classification regression tree CART, random forest RF, gradient boosting machine GBM, support vector machine SVM, artificial neural network ANN; using random search method for all relationship models to cross-validate the correlation coefficient Q 2 CV The optimal hyperparameters of the model are determined by using the 8. The biological toxicity prediction method based on SMOTE small sample machine learning according to claim 1, characterized in that: The step 5 includes the following formula: where Q 2 CV represents the cross-validation correlation coefficient, Q 2 ext represents the external validation coefficient of determination, the subscripts CV and ext represent cross validation and external validation, respectively, RMSE CV或ext Represents the cross-validation or external validation root mean square error, MAE CV represents the cross-validation mean absolute error, R 2 Represents the correlation coefficient of the training set, y i Indicates the acute toxicity observation value, Indicates the acute toxicity prediction value, represents the average value of all acute toxicity observations, i represents the sample number, and n represents the total number of samples.