Microplastic toxicity prediction method fusing machine learning and meta-analysis
By integrating machine learning and meta-analysis, this study performs multi-dimensional detection and data quality assessment of microplastic samples, establishes dynamic weight allocation and biological constraints, solves the problems of data heterogeneity and multi-toxicity endpoint prediction in microplastic toxicity detection, and achieves highly accurate and reliable toxicity prediction and risk assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies for microplastic toxicity detection suffer from data heterogeneity, lack of biological constraints in predicting multiple toxicity endpoints, and insufficient assessment of the reliability of prediction results.
We employ a method that combines machine learning and meta-analysis to detect the spectra, morphology, and surface charge of microplastic samples. We establish a dynamic weight allocation mechanism to assess data quality and determine biological constraints. We train a constraint fusion prediction algorithm through collaborative learning and combine it with Bayesian quantification to output risk assessment.
It improves the accuracy and reliability of microplastic toxicity prediction, achieves biological rationality and predictive stability among multiple toxic endpoints, has real-time early warning capabilities, and provides quantitative risk assessment results.
Smart Images

Figure CN120951142B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of toxicity detection technology, and in particular to a method for predicting the toxicity of microplastics that integrates machine learning and meta-analysis. Background Technology
[0002] Microplastic toxicity assessment is an important research direction in the field of environmental science. Current technologies mainly employ two methods: laboratory biotoxicity testing and computer-aided prediction. Laboratory testing methods use model organisms such as zebrafish, algae, and bacteria to detect toxicity endpoints such as LC50 and EC50. Computer prediction methods are based on quantitative structure-activity relationship (QSAR) models, using chemical structure parameters to predict toxic effects. Meta-analysis, as a statistical method, is used to integrate the results of multiple independent studies, while machine learning algorithms are also widely used in the field of toxicity prediction.
[0003] However, existing technologies have significant shortcomings: First, different studies use different detection methods, experimental conditions, and quality control standards, resulting in serious data heterogeneity problems, and simple data merging cannot effectively eliminate systematic bias; second, traditional meta-analysis only assigns weights based on sample size, ignoring the impact of the standardization of detection methods on data quality and lacking intelligent assessment of data reliability; third, existing machine learning methods usually predict each toxicity endpoint independently, failing to fully utilize the biological associations between different toxicity mechanisms. Summary of the Invention
[0004] This application provides a microplastic toxicity prediction method that integrates machine learning and meta-analysis, which solves the problems of inaccurate correction for heterogeneity of detection data, lack of biological constraints in prediction of multiple toxic endpoints, and insufficient reliability assessment of prediction results, thereby improving the accuracy and reliability of microplastic toxicity prediction.
[0005] This application provides a method for predicting microplastic toxicity by integrating machine learning and meta-analysis. The method includes:
[0006] The physicochemical characteristic parameters of microplastics were obtained by performing spectral analysis, morphology analysis, and surface charge analysis on the microplastic samples.
[0007] The physicochemical characteristic parameters of the microplastics and the biotoxicity detection data were meta-analyzed and processed. A dynamic weight allocation was established based on the standardized scoring of the detection methods to obtain quality-corrected toxicity data.
[0008] Determine whether the correlation between acute toxicity and chronic toxicity in the quality-corrected toxicity data meets the biological constraints. If it does, perform collaborative learning training based on the constraints between multiple toxic endpoints to establish the association mapping between physicochemical characteristics and multiple toxic endpoints, and obtain the constraint fusion prediction algorithm.
[0009] The physicochemical characteristic parameters of the sample to be tested are input into the constrained fusion prediction algorithm to calculate toxicity. It is determined whether the prediction uncertainty exceeds the set threshold. If it does not exceed the threshold, the toxicity prediction value and confidence interval are output in combination with Bayesian quantization.
[0010] Based on the predicted toxicity value, the risk quotient is calculated and the mechanism contribution is analyzed to determine whether the risk quotient exceeds the safety threshold. If it does, it is marked as high risk and an early warning message is generated, thus obtaining the graded risk assessment result.
[0011] Optionally, a Fourier transform infrared spectroscopy is used to perform full-band spectral scanning on the microplastic sample. The matching degree with the standard polymer spectrum is calculated based on the characteristic peak intensity ratio to obtain chemical composition detection data. The microplastic sample is then examined by scanning electron microscopy, and the particle outline is automatically identified and geometric parameters are measured. The equivalent circle diameter and aspect ratio are calculated to obtain morphology detection data. The electrophoretic mobility of the microplastic sample under various pH conditions is measured based on dynamic light scattering, and the Zeta potential distribution is calculated according to the Henry's function formula to obtain surface charge detection data. The chemical composition detection data, morphology detection data, and surface charge detection data are integrated according to a standardized format to obtain the physicochemical characteristic parameters of the microplastic.
[0012] Optionally, a data quality assessment is performed on the microplastic physicochemical characteristic parameters and multi-source biotoxicity detection data. A reliability score for each data source is calculated based on the experimental design integrity score and detection accuracy score, resulting in a data quality assessment matrix. Weights are calculated based on the reliability scores and sample size information in the data quality assessment matrix. The reliability scores are multiplied by the normalized sample size weights to obtain dynamic weight allocation coefficients. Based on these dynamic weight allocation coefficients, toxicity detection data from different sources are weighted and merged. A DerSimonian-Laird random effects model is used to calculate the weighted average effect size, yielding heterogeneity correction results. These heterogeneity correction results are then fused with the microplastic physicochemical characteristic parameters and categorized according to toxicity endpoint type to obtain the quality-corrected toxicity data.
[0013] Optionally, Pearson correlation is calculated on the acute toxicity data and chronic toxicity data in the quality-corrected toxicity data to determine whether the correlation coefficient is greater than a preset biological constraint threshold. If it is greater, the biological constraint condition is confirmed to be met, and the toxicity endpoint correlation verification result is obtained. A multi-task learning loss function is established based on the toxicity endpoint correlation verification result, and the loss terms of acute toxicity, chronic toxicity, bioaccumulation, and genotoxicity are weighted and combined to obtain a constraint loss function. A neural network is trained on the physicochemical characteristic parameters of the microplastics and the toxicity endpoint data based on the constraint loss function, and a shared hidden layer is used to extract general feature representations to obtain a multi-toxicity endpoint mapping network. The multi-toxicity endpoint mapping network is integrated and optimized with the biological constraint condition, and a regularization constraint term between toxicity endpoints is added to obtain the constraint fusion prediction algorithm.
[0014] Optionally, the acute toxicity LC50 value and chronic toxicity NOEC value in the quality-corrected toxicity data are logarithmically transformed to convert the values into a normal distribution form, resulting in standardized toxicity data pairs. Based on these standardized toxicity data pairs, the Pearson correlation coefficient is calculated, and the correlation is quantified using the formula covariance divided by the standard deviation product, yielding the toxicity endpoint correlation coefficient. This toxicity endpoint correlation coefficient is then compared with the biological constraint threshold of 0.6 to determine whether acute and chronic toxicity meet the positive correlation requirement, thus obtaining the constraint condition judgment result. Based on the constraint condition judgment result, it is determined whether to perform multi-toxicity endpoint collaborative learning. If the constraint condition is met, the data is marked as qualified and enters the subsequent training process, resulting in the toxicity endpoint correlation verification result.
[0015] Optionally, the physicochemical characteristic parameters of the sample to be tested are input into the constrained fusion prediction algorithm for forward propagation calculation. The predicted values of acute toxicity, chronic toxicity, bioaccumulation, and genotoxicity are calculated through a shared hidden layer and a task-specific output layer, respectively, to obtain multi-terminal toxicity prediction data. Based on Monte Carlo Dropout sampling, the multi-terminal toxicity prediction data is forward propagated multiple times to calculate the prediction variance and cognitive uncertainty, thereby obtaining a prediction uncertainty quantification result. The prediction uncertainty quantification result is numerically compared with a preset uncertainty threshold to determine whether the total uncertainty exceeds the set threshold. If it does not exceed the threshold, the prediction result is confirmed to be reliable, and a reliability judgment result is obtained. Based on the reliability judgment result and Bayesian posterior distribution calculation, the confidence interval of the multi-terminal toxicity prediction value is quantified to obtain a prediction result containing the toxicity prediction value and the confidence interval.
[0016] Optionally, a risk quotient is calculated based on the predicted toxicity value and the predicted no-effect concentration. The predicted environmental concentration is divided by the predicted no-effect concentration to obtain the risk quotient data. The risk level is classified according to the magnitude of the risk quotient to obtain a quantitative risk assessment result. SHAP value analysis is performed on the predicted toxicity value to quantify the contribution of each physicochemical characteristic parameter to the toxicity prediction result, identify the dominant toxicity mechanism and key influencing factors, and obtain the mechanism contribution analysis result. The risk quotient in the quantitative risk assessment result is compared with the safety threshold of 1.0 to determine whether the risk quotient exceeds the safety threshold. If it does, it is marked as a high-risk level and an early warning mechanism is triggered, resulting in a risk level determination result. A standardized assessment report is generated based on the risk level determination result and the mechanism contribution analysis result, integrating the risk level, toxicity mechanism explanation, and management recommendations to obtain the graded risk assessment result.
[0017] The technical solution provided in this application combines multi-dimensional detection technologies, including spectral detection, morphology detection, and surface charge detection, to comprehensively acquire key physicochemical characteristic parameters such as chemical composition, geometric morphology, and surface properties of microplastics, providing a complete characteristic basis for subsequent toxicity prediction. When performing meta-analysis of microplastic physicochemical characteristic parameters and biotoxicity detection data, a dynamic weight allocation mechanism based on standardized scoring of detection methods effectively solves the data heterogeneity problem in traditional meta-analysis. By intelligently assessing the reliability of different data sources and dynamically adjusting weights, the accuracy and consistency of quality-corrected toxicity data are significantly improved. The design of biological constraints for judging the correlation between acute and chronic toxicity in quality-corrected toxicity data ensures the rationality of the constraint relationships between multiple toxicity endpoints. The correlation mapping between physicochemical characteristics and multiple toxicity endpoints established through collaborative learning training fully utilizes the intrinsic connections between different toxicity mechanisms, making the constraint fusion prediction algorithm more biologically reasonable and predictively stable.
[0018] In the process of inputting the physicochemical characteristic parameters of the sample to be tested into the constrained fusion prediction algorithm for toxicity calculation, the prediction uncertainty threshold judgment mechanism effectively ensures the reliability of the prediction results. Combined with the Bayesian quantification output of the toxicity prediction value and confidence interval, it provides quantitative credibility information for risk assessment. The design of calculating the risk quotient and analyzing the mechanism contribution based on the toxicity prediction value not only achieves quantitative assessment of the risk level but also enhances the interpretability of the prediction results through mechanism explanation. The intelligent judgment function that automatically marks it as high-risk and generates early warning information when the risk quotient exceeds the safety threshold enables the graded risk assessment results to have real-time early warning capabilities. In particular, in the specific application field of microplastic toxicity prediction, the multi-task collaborative learning algorithm fully considers the biological correlations between multiple endpoints such as acute toxicity, chronic toxicity, bioaccumulation, and genotoxicity. The design of the constrained fusion prediction algorithm avoids biologically unreasonable results that may occur when predicting each endpoint individually, significantly improving prediction accuracy and the biological credibility of the prediction results. Attached Figure Description
[0019] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a schematic diagram of an embodiment of the microplastic toxicity prediction method that integrates machine learning and meta-analysis in this application.
[0021] Figure 2 This is a flowchart of the microplastic toxicity data quality correction and fusion analysis in the embodiments of this application;
[0022] Figure 3 This is a graph showing the correlation analysis and data verification of acute and chronic toxicity of polystyrene microplastics in the embodiments of this application. Detailed Implementation
[0023] This application provides a method for predicting microplastic toxicity by integrating machine learning and meta-analysis. The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data used can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" or "having" and any variations thereof are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0024] For ease of understanding, the specific process of the embodiments of this application is described below. Please refer to [link / reference]. Figure 1 One embodiment of the microplastic toxicity prediction method integrating machine learning and meta-analysis in this application includes:
[0025] Step S1: Perform spectral detection, morphology detection, and surface charge detection on the microplastic sample to obtain the physicochemical characteristic parameters of the microplastic.
[0026] Specifically, when testing microplastic samples, a full-band spectral scan is performed using a Fourier transform infrared spectrometer. The chemical composition characteristics of the sample are determined by analyzing the matching degree between the intensity ratio of characteristic peaks and the standard polymer spectrum, thereby obtaining spectral detection data. The sample is then characterized under a scanning electron microscope. By automatically identifying particle outlines and measuring geometric parameters, the equivalent circle diameter and aspect ratio are calculated to obtain morphological detection data reflecting the particle shape and size distribution. Based on this, the electrophoretic mobility of the sample under different pH conditions is monitored using dynamic light scattering methods, and the distribution of Zeta potential is calculated according to the Henry's function formula, thereby obtaining surface charge detection data. Through the comprehensive detection results of the above three dimensions of spectrum, morphology, and surface charge, the chemical composition, appearance structure characteristics, and surface charge state of microplastics can be comprehensively described. In the data processing stage, these detection data are uniformly converted into a standardized format and effectively integrated to form complete physicochemical characteristic parameters of microplastics that can be used for modeling and analysis.
[0027] Step S2: Meta-analysis of the physicochemical characteristic parameters of microplastics and biotoxicity detection data is performed, and dynamic weight allocation is established based on the standardized scoring of detection methods to obtain quality-corrected toxicity data.
[0028] Specifically, the obtained physicochemical characteristics of microplastics are systematically integrated with biotoxicity detection data from different sources. This requires a quality assessment of the reliability of each data source. Data is quantitatively scored using experimental design integrity and detection accuracy scores to form a data quality assessment matrix, reflecting the credibility levels of different research results. Based on this, the credibility score is combined with sample size information, and after normalization, a dynamic weighting coefficient is calculated. This ensures that high-quality data with larger sample sizes receive greater weight in subsequent analyses, while the impact of low-quality or insufficient sample sizes is relatively weakened. The DerSimonian-Laird random effects model is used to calculate effect sizes and correct for heterogeneity in the weighted toxicity data to eliminate systematic biases caused by differences in experimental conditions and detection methods. The heterogeneously corrected toxicity results are then fused and organized with the previously obtained physicochemical characteristics of microplastics, and categorized according to different toxicity endpoint types to obtain representative quality-corrected toxicity data.
[0029] Step S3: Determine whether the correlation between acute toxicity and chronic toxicity in the quality-corrected toxicity data meets the biological constraints. If it does, perform collaborative learning training based on the constraints between multiple toxic endpoints to establish the association mapping between physicochemical characteristics and multiple toxic endpoints, and obtain the constraint fusion prediction algorithm.
[0030] Specifically, in the process of correlation determination and collaborative learning, the acute and chronic toxicities in the quality-corrected toxicity data are standardized to ensure comparability of the data at the same scale. Correlation analysis is used to examine the relationship between the two types of toxicity endpoints and to determine whether they meet the predetermined biological constraints. If the determination results show that the correlation meets the requirements, the dataset is further introduced into a multi-task collaborative learning framework. In this framework, a general representation of physicochemical features is extracted through a shared hidden layer, and multiple endpoints such as acute toxicity, chronic toxicity, bioaccumulation, and genotoxicity are respectively represented in the output layer. By using the inter-endpoint constraint term and regularization condition introduced in the loss function, the consistency and rationality of different toxicity prediction results at the biological level are ensured. During the training process, dynamic weight adjustment and model optimization mechanisms are combined to enable multi-endpoint predictions to complement each other on the basis of information sharing, establish the correlation mapping between physicochemical features and multiple toxicity endpoints, and obtain a constraint fusion prediction algorithm that combines prediction accuracy and biological credibility.
[0031] Step S4: Input the physicochemical characteristic parameters of the sample to be tested into the constraint fusion prediction algorithm to calculate toxicity, determine whether the prediction uncertainty exceeds the set threshold, and if it does not exceed the threshold, output the toxicity prediction value and confidence interval by combining Bayesian quantization.
[0032] Specifically, in toxicity prediction, the physicochemical characteristics of the sample to be tested are input into the established constrained fusion prediction algorithm. Common features are extracted through a shared hidden layer, and multi-endpoint prediction results such as acute toxicity, chronic toxicity, bioaccumulation, and genotoxicity are generated in different task output layers to obtain complete toxicity prediction data. To avoid the uncertainty of a single prediction result affecting the overall reliability, a random sampling method is used to perform multiple forward propagations on the model to calculate the variance and uncertainty level of the prediction results. The results are then compared with a preset threshold to determine whether the prediction has sufficient credibility. If the uncertainty is within an acceptable range, the predicted value is further quantified by combining the Bayesian posterior distribution to generate a corresponding confidence interval for each toxicity endpoint, thereby outputting a result that includes both the specific predicted value and the confidence range.
[0033] Step S5: Calculate the risk quotient and analyze the mechanism contribution based on the toxicity prediction value. Determine whether the risk quotient exceeds the safety threshold. If it does, mark it as high risk and generate early warning information to obtain the graded risk assessment result.
[0034] Specifically, after completing the toxicity prediction, the predicted toxicity endpoint values are compared with potential concentration information in the environment. A risk quotient is calculated to characterize the relative relationship between receptor exposure levels and no-effect levels, thus obtaining quantitative risk assessment data. Interpretive analysis methods are used to decompose the prediction results, and feature importance assessment methods are employed to identify the contribution of different physicochemical parameters to the toxicity prediction, revealing potential mechanisms of action and key influencing factors, and providing interpretability support for the results. Based on this, the risk quotient is compared with a set safety threshold. If the result exceeds the safety threshold, it is automatically marked as a high-risk level, triggering an early warning mechanism to indicate potential ecological and health hazards. The risk level determination results are integrated with the mechanism contribution analysis results to generate a standardized assessment report containing risk level classification, explanation of toxicity mechanisms, and corresponding management recommendations, thus obtaining a complete graded risk assessment result.
[0035] It is understood that the executing entity of this application can be a microplastic toxicity prediction system integrating machine learning and meta-analysis, or it can be a terminal or a server; the specific implementation is not limited here. This application's embodiments use a server as an example for illustration.
[0036] In one specific embodiment, the process of performing step S1 may specifically include the following steps:
[0037] (1) The microplastic sample was scanned across the entire spectrum using a Fourier transform infrared spectrometer. The matching degree between the characteristic peak intensity ratio and the standard polymer spectrum was calculated to obtain the chemical composition detection data.
[0038] (2) The microplastic sample is examined by scanning electron microscopy. The particle outline is automatically identified and the geometric parameters are measured. The equivalent circle diameter and aspect ratio are calculated to obtain the morphology detection data.
[0039] (3) Based on dynamic light scattering, the electrophoretic mobility of microplastic samples under various pH conditions was measured, and the Zeta potential distribution was calculated according to the Henry function formula to obtain surface charge detection data;
[0040] (4) The chemical composition detection data, morphology detection data and surface charge detection data are integrated in a standardized format to obtain the physicochemical characteristic parameters of microplastics.
[0041] Specifically, a full-band spectral scan of the microplastic samples was performed using Fourier transform infrared spectroscopy to obtain chemical composition information. During the scan, the matching degree with the standard polymer spectrum was calculated by combining the characteristic peak intensity ratio, thereby determining the chemical composition of the sample and generating chemical composition detection data. The microplastic samples were then examined under a scanning electron microscope. By automatically identifying the particle outline and measuring geometric parameters, the equivalent circle diameter and aspect ratio of the sample were calculated, thus obtaining detailed data on particle morphology. This process provided crucial information for morphology detection. Electrophoretic mobility measurements of the microplastic samples were performed using dynamic light scattering technology to determine their electrophoretic behavior under different pH conditions. The Zeta potential distribution of the samples was calculated based on the Henry's function formula. This analysis revealed the surface charge characteristics of the samples, obtaining surface charge detection data. Through these analytical steps, the physicochemical properties of the microplastic samples, including their chemical composition, morphological characteristics, and surface charge characteristics, were comprehensively characterized, ensuring accurate and comprehensive basic data for subsequent toxicity prediction. All test results, including chemical composition data, morphology data, and surface charge data, are integrated in a unified standard format to form the physicochemical characteristic parameters of microplastics.
[0042] For example, when performing a full-spectral scan of a microplastic sample using Fourier transform infrared spectroscopy (FTIR), assuming a common polyethylene (PE) microplastic sample was selected, the spectral scan results showed a distinct CH stretching vibration peak near approximately 2900 cm⁻¹ and a C=O stretching vibration peak near 1700 cm⁻¹. The intensity ratios of these characteristic peaks showed a high degree of match with the spectrum of standard polyethylene, thus confirming that the sample's chemical composition was polyethylene. The resulting chemical composition data provided a clear basis for material identification in subsequent toxicity assessments.
[0043] For example, in surface charge detection, when using dynamic light scattering (DLS) to measure the zeta potential of a microplastic sample, assuming an analysis of a polystyrene (PS) sample at pH 7, a zeta potential of -30 mV was obtained. This potential value indicates that the microplastic particles have a strong negative charge in aqueous solution, which is important for assessing their interactions with other chemicals and organisms in the water.
[0044] In one specific embodiment, the process of performing step S2 may specifically include the following steps:
[0045] (1) Data quality assessment was conducted on the physicochemical characteristic parameters of microplastics and the multi-source biotoxicity detection data. The credibility score of each data source was calculated based on the experimental design integrity score and the detection accuracy score, and the data quality assessment matrix was obtained.
[0046] (2) Calculate the weights based on the credibility score and sample size information in the data quality assessment matrix, and multiply the credibility score by the normalized weight of the sample size to obtain the dynamic weight allocation coefficient.
[0047] (3) The toxicity detection data from different sources were weighted and merged based on the dynamic weight allocation coefficient, and the weighted average effect size was calculated using the DerSimonian-Laird random effects model to obtain the heterogeneity correction results.
[0048] (4) The heterogeneity correction results are fused with the physicochemical characteristic parameters of microplastics, and the data are classified and organized according to the type of toxicity endpoint to obtain quality-corrected toxicity data.
[0049] Specifically, when assessing the data quality of microplastic physicochemical characteristics and multi-source biotoxicity detection data, an experimental design integrity scoring system was first established. This scoring system includes three dimensions: the degree of standardization in sample selection, the consistency of experimental condition control, and the adequacy of repeated experiments. The degree of standardization in sample selection assesses whether the data source uses a unified microplastic sample preparation method and particle size screening standard. The consistency of experimental condition control assesses the accuracy of controlling environmental parameters such as temperature, pH, and ionic strength. The adequacy of repeated experiments assesses whether the number of repeated measurements for each toxicity endpoint meets statistical requirements. Each dimension is scored on a ten-point scale, and the scores of the three dimensions are weighted and averaged using weighting coefficients of 0.4, 0.3, and 0.3 to calculate the experimental design integrity score. The detection accuracy score is quantitatively evaluated based on three technical indicators: precision, accuracy, and limit of detection (LOD). Precision is reflected by the relative standard deviation (RSD), accuracy by the recovery rate, and the LOD by the method sensitivity parameter. Each technical indicator is also scored on a ten-point scale, and the detection accuracy score is calculated using weighting coefficients of 0.5, 0.3, and 0.2. The experimental design integrity score and the detection accuracy score are weighted and calculated with weighting coefficients of 0.6 and 0.4 respectively to obtain the credibility score of each data source. The credibility scores of all data sources constitute the data quality assessment matrix.
[0050] When calculating weights based on the credibility scores and sample size information in the data quality assessment matrix, the credibility scores of each data source are first normalized to their maximum and minimum values, standardizing the score range to the interval between 0 and 1. Simultaneously, the sample size data undergoes a logarithmic transformation followed by normalization to eliminate the impact of excessively large differences in sample size values on weight calculation. The normalized credibility scores and normalized sample size data are then linearly combined using weighting coefficients of 0.7 and 0.3 to calculate the comprehensive weight score for each data source. The comprehensive weight scores of all data sources are then normalized so that the sum of all weights equals 1, yielding the dynamic weight allocation coefficient. This dynamic weight allocation coefficient reflects both the quality level and sample size of the data source; data sources with high quality and large sample sizes receive greater weights, while data sources with low quality or insufficient sample sizes receive relatively smaller weights.
[0051] When weighting and merging toxicity testing data from different sources using dynamic weighting coefficients, the DerSimonian-Laird random effects model is employed for effect size estimation. This model first calculates the effect size and variance within each data source. The effect size is expressed as the standardized mean difference (SMD) or log-odds ratio (LRR), and the variance is calculated based on the sample size and standard deviation. Then, the effect sizes from each data source are weighted and averaged using the dynamic weighting coefficients to calculate the weighted effect size. The DerSimonian-Laird model corrects for systematic differences between data sources by estimating the heterogeneity variance τ² between studies. This heterogeneity variance is calculated using the Q statistic and degrees of freedom. The final weighted average effect size, combined with the heterogeneity correction parameters, yields a heterogeneity-corrected result that eliminates systematic bias. This heterogeneity-corrected result effectively integrates toxicity data from different laboratories and testing methods, eliminating data bias caused by differences in experimental conditions and methods.
[0052] When fusing heterogeneity correction results with microplastic physicochemical characteristic parameters, the results are first categorized according to the type of toxicity endpoint, including different types such as acute toxicity LC50 value, chronic toxicity NOEC value, bioaccumulation BCF value, and genotoxic DNA damage index. A correspondence is established between the correction result for each toxicity endpoint and the corresponding microplastic physicochemical characteristic parameters, including key parameters such as chemical composition type, particle size distribution, surface charge density, and specific surface area. The data fusion process uses a standardized data format, uniformly converting physicochemical parameters and toxicity data of different dimensions into dimensionless standardized values. The fused dataset is indexed in a three-level classification system according to microplastic material type, particle size range, and toxicity endpoint type, forming a structured quality-corrected toxicity dataset, which is directly used as input data for subsequent machine learning model training.
[0053] refer to Figure 2 The figure illustrates the flowchart for quality correction and fusion analysis of microplastic toxicity data. This method yields quality-corrected toxicity data that not only considers the differences in quality and sample size across various data sources but also accurately reflects the comprehensiveness and consistency of the toxicity data.
[0054] When weighting and merging toxicity test data from different sources, assuming there are three sets of toxicity data from laboratories A, B, and C, after weighting, the data from laboratory A has a weight of 0.5, the data from laboratory B has a weight of 0.3, and the data from laboratory C has a weight of 0.2. Using the DerSimonian-Laird random effects model, a comprehensive effect size can be calculated based on the weighted data. During data fusion, it is assumed that data on different toxicity endpoints, such as acute toxicity, chronic toxicity, bioaccumulation, and genotoxicity, are extracted from the quality-corrected toxicity data. By classifying and organizing this data according to different toxicity endpoint types, more accurate and representative quality-corrected toxicity data can be obtained.
[0055] In one specific embodiment, the process of performing step S3 may specifically include the following steps:
[0056] (1) Perform Pearson correlation calculation on acute toxicity data and chronic toxicity data in quality-corrected toxicity data, and determine whether the correlation coefficient is greater than the preset biological constraint threshold. If it is greater, it is confirmed that the biological constraint conditions are met, and the toxicity endpoint correlation verification results are obtained.
[0057] (2) Based on the results of the correlation verification of toxicity endpoints, a multi-task learning loss function is established, and the loss terms of acute toxicity, chronic toxicity, bioaccumulation and genotoxicity are weighted and combined to obtain the constraint loss function.
[0058] (3) Based on the constraint loss function, a neural network is trained on the physicochemical characteristic parameters and toxicity endpoint data of microplastics. A shared hidden layer is used to extract general feature representations to obtain a multi-toxicity endpoint mapping network.
[0059] (4) The multi-toxic endpoint mapping network is integrated and optimized with biological constraints, and regularization constraints between toxic endpoints are added to obtain the constraint fusion prediction algorithm.
[0060] Specifically, Pearson correlation calculations were performed on the acute and chronic toxicity data in the quality-corrected toxicity data. The correlation coefficient between the two was calculated through statistical analysis. If the coefficient exceeded a preset biological constraint threshold, a significant positive correlation between acute and chronic toxicity was confirmed, satisfying the biological constraint conditions. The correlation verification results of the toxicity endpoints were obtained, confirming the biological consistency and correlation among these data. Based on the correlation verification results, a multi-task learning loss function was established, weighting the loss terms for acute toxicity, chronic toxicity, bioaccumulation, and genotoxicity to ensure a reasonable balance and optimization of the learning objectives for different toxicity endpoints during training, thus constructing a constraint loss function encompassing all toxicity endpoints. Based on this constraint loss function, a neural network was used to train the physicochemical characteristics of microplastics and the toxicity endpoint data, extracting general feature representations through shared hidden layers. This process improved the network's learning ability across multiple tasks and ensured the universality and sharing of feature extraction, enabling the neural network to learn the complex relationship between the physicochemical characteristics of microplastics and multiple toxicity endpoints. The trained multi-toxic endpoint mapping network is integrated and optimized with biological constraints, and regularization constraints between toxic endpoints are further added to make the learning process of each toxic endpoint more stable and reasonable, resulting in a constraint fusion prediction algorithm.
[0061] In one specific embodiment, the process of performing Pearson correlation calculation on acute toxicity data and chronic toxicity data in quality-corrected toxicity data, determining whether the correlation coefficient is greater than a preset biological constraint threshold, and confirming that the biological constraint condition is met to obtain the toxicity endpoint correlation verification result can specifically include the following steps:
[0062] (1) Logarithmic transformation was performed on the acute toxicity LC50 value and chronic toxicity NOEC value in the quality-corrected toxicity data to convert the values into a normal distribution form, thus obtaining standardized toxicity data pairs;
[0063] (2) Based on standardized toxicity data, the Pearson correlation coefficient was calculated, and the correlation was quantified by the formula of covariance divided by standard deviation to obtain the correlation coefficient of toxicity endpoint;
[0064] (3) Compare the correlation coefficient of the toxicity endpoint with the biological constraint threshold of 0.6 to determine whether the acute toxicity and chronic toxicity meet the positive correlation requirement and obtain the constraint condition judgment result.
[0065] (4) Determine whether to perform multi-toxic endpoint collaborative learning based on the constraint judgment results. If the constraint conditions are met, mark it as qualified data and enter the subsequent training process to obtain the toxic endpoint correlation verification results.
[0066] Specifically, the acute toxicity LC50 and chronic toxicity NOEC values in the quality-corrected toxicity data were logarithmically transformed to convert them into a normal distribution, thereby eliminating data skewness and ensuring that the data met the requirements of subsequent statistical analysis, resulting in standardized toxicity data pairs. Based on this, Pearson correlation was calculated to assess the correlation between acute and chronic toxicity based on the standardized toxicity data. The calculation formula is as follows: ,in, The Pearson correlation coefficient is used. and The data are standardized LC50 for acute toxicity and NOEC for chronic toxicity, respectively. and The mean and covariance of acute toxicity and chronic toxicity are respectively. This reflects the linear relationship between the two datasets, and the square root of the denominator is the product of their respective standard deviations, ensuring the quantification of the correlation coefficient. The calculated correlation coefficient of the toxicity endpoint is compared numerically with a preset biological constraint threshold of 0.6. If the correlation coefficient is greater than this threshold, it indicates a significant positive correlation between acute and chronic toxicity, satisfying the biological constraint. Based on this constraint, it is determined whether to proceed to the multi-toxicity endpoint collaborative learning process. If the condition is met, the data is marked as qualified and enters the subsequent training process to obtain the toxicity endpoint correlation verification results.
[0067] Taking polystyrene (PS) microplastics as an example, assuming the acute toxicity LC50 value is 15 mg / L and the chronic toxicity NOEC value is 5 mg / L measured in the experiment, a logarithmic transformation was performed on these two values to ensure the data conforms to a normal distribution, resulting in logarithmic values of 1.176 and 0.699, respectively. The Pearson correlation coefficient was calculated using these two standardized toxicity data, yielding a correlation coefficient r of 0.82, indicating a strong positive correlation between acute and chronic toxicity. Comparing this result with the preset biological constraint threshold of 0.6, the r value was found to be greater than the threshold, meeting the biological constraint conditions. The correlation between acute and chronic toxicity was confirmed to be satisfactory, and the data was marked as qualified. These qualified data were then input into a multi-toxicity endpoint collaborative learning model for subsequent training, yielding the toxicity endpoint correlation verification results for this microplastic material. Figure 3 This figure shows the correlation analysis and data validation of acute and chronic toxicity of polystyrene microplastics.
[0068] In one specific embodiment, the process of performing step S4 may specifically include the following steps:
[0069] (1) Input the physicochemical characteristic parameters of the sample to be tested into the constraint fusion prediction algorithm for forward propagation calculation. The predicted values of acute toxicity, chronic toxicity, bioaccumulation and genotoxicity are calculated by the shared hidden layer and the task-specific output layer, respectively, to obtain multi-endpoint toxicity prediction data.
[0070] (2) Based on Monte Carlo Dropout sampling, multiple forward propagations were performed on the multi-terminal toxicity prediction data to calculate the prediction variance and cognitive uncertainty, and the prediction uncertainty quantification results were obtained.
[0071] (3) Compare the quantified result of the prediction uncertainty with the preset uncertainty threshold to determine whether the total uncertainty exceeds the set threshold. If it does not exceed the threshold, the prediction result is confirmed to be reliable and the reliability judgment result is obtained.
[0072] (4) Based on the reliability judgment results and the Bayesian posterior distribution calculation, the confidence interval of the multi-terminal toxicity prediction value is quantified to obtain the prediction result containing the toxicity prediction value and the confidence interval.
[0073] Specifically, taking a certain polymer microplastic as an example, the physicochemical characteristic parameters of the sample to be tested are input into a constrained fusion prediction algorithm for forward propagation calculation. The algorithm extracts general features through a shared hidden layer, and then calculates the predicted values of acute toxicity, chronic toxicity, bioaccumulation, and genotoxicity through a task-specific output layer. This process yields multi-terminal toxicity prediction data. To quantify the uncertainty of the prediction, the multi-terminal toxicity prediction data is forward-propagated multiple times using the Monte Carlo Dropout sampling method, each time with a different Dropout randomness, thereby calculating the prediction variance and cognitive uncertainty, and obtaining the quantified prediction uncertainty result. The quantified prediction uncertainty result is compared with a preset uncertainty threshold. If the quantified result does not exceed the threshold, the reliability of the prediction result is confirmed, and a reliability judgment result is obtained. Combining the Bayesian posterior distribution, the confidence interval of the multi-terminal toxicity prediction values is quantified, thus obtaining a prediction result containing the toxicity prediction value and its confidence interval.
[0074] For example, in Monte Carlo Dropout sampling, suppose the above prediction data is sampled 1000 times. The Dropout mechanism randomly discards a portion of neurons in the neural network, generating 1000 different predictions. By calculating the variance of these results, we assume the variance for acute toxicity is 0.2, for chronic toxicity is 0.1, for bioaccumulation is 0.05, and for genotoxicity is 0.04. This process helps quantify the cognitive uncertainty of the prediction results.
[0075] For example, when determining whether the prediction uncertainty exceeds a set threshold, let's assume the preset uncertainty threshold is 0.15. When comparing the calculated variance with this threshold, the variances of acute toxicity and chronic toxicity are 0.2 and 0.1, respectively, neither exceeding the threshold. Therefore, the prediction result can be confirmed as reliable, and a judgment result can be obtained.
[0076] In one specific embodiment, the process of performing step S5 may specifically include the following steps:
[0077] (1) The risk quotient is calculated based on the predicted toxicity value and the predicted no-effect concentration. The predicted environmental concentration is divided by the predicted no-effect concentration to obtain the risk quotient data. The risk level is classified according to the size of the risk quotient to obtain the quantitative risk assessment result.
[0078] (2) SHAP value analysis was performed on the toxicity prediction values to quantify the contribution of each physicochemical characteristic parameter to the toxicity prediction results, identify the dominant toxicity mechanism and key influencing factors, and obtain the mechanism contribution analysis results.
[0079] (3) Compare the risk quotient in the quantitative risk assessment results with the safety threshold of 1.0 to determine whether the risk quotient exceeds the safety threshold. If it does, mark it as a high-risk level and trigger the early warning mechanism to obtain the risk level determination result.
[0080] (4) Generate a standardized assessment report based on the risk level determination results and mechanism contribution analysis results, integrate the risk level, toxicity mechanism explanation and management recommendations to obtain the graded risk assessment results.
[0081] Specifically, a risk quotient is calculated based on the predicted toxicity value and the predicted no-effect concentration. The predicted environmental concentration is divided by the predicted no-effect concentration to obtain the risk quotient data. Based on the magnitude of the risk quotient, risk levels are further classified to obtain a quantitative risk assessment result. SHAP value analysis is performed on the predicted toxicity values. By calculating the contribution of each physicochemical characteristic parameter to the toxicity prediction result, the dominant toxicity mechanism and key influencing factors are identified, resulting in a mechanism contribution analysis. This clarifies which physicochemical characteristics play an important role in toxicity prediction and reveals the main toxicity mechanisms. The risk quotient in the quantitative risk assessment result is numerically compared with a safety threshold to determine whether the risk quotient exceeds the set safety threshold. If it does, it is marked as a high-risk level and an early warning mechanism is triggered, resulting in a risk level determination result. A standardized assessment report is generated based on the risk level determination result and the mechanism contribution analysis result, integrating the risk level, toxicity mechanism explanation, and management recommendations to obtain a graded risk assessment result.
[0082] For example, when calculating the risk quotient, if the predicted concentration level of a microplastic sample in the environment is higher than its ineffective concentration, the calculated risk quotient will be greater than 1, indicating that the sample may pose a significant risk to the ecological environment, and the risk level can be classified as high. During the risk level determination process, after comparing the risk quotient with a safety threshold, if the result shows that the risk quotient is significantly higher than the threshold, the system will mark the sample as high-risk and automatically trigger an early warning mechanism to indicate potential environmental hazards. By integrating the quantitative results of the risk level with the mechanism contribution analysis results, the report not only includes the risk level classification but also details the explanation of the toxicity mechanism and corresponding management recommendations, thus making the assessment results scientific, interpretable, and practical.
[0083] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for predicting microplastic toxicity by integrating machine learning and meta-analysis, characterized in that, The method includes: Step S1: Perform spectral analysis, morphology analysis, and surface charge analysis on the microplastic sample to obtain the physicochemical characteristic parameters of the microplastic; Step S2: Meta-analysis of the microplastic physicochemical characteristic parameters and biotoxicity detection data is performed. A dynamic weight allocation is established based on the standardized scoring of the detection methods to obtain quality-corrected toxicity data. This includes: assessing the data quality of the microplastic physicochemical characteristic parameters and multi-source biotoxicity detection data; calculating the reliability score of each data source based on experimental design integrity and detection accuracy scores to obtain a data quality assessment matrix; calculating weights based on the reliability scores and sample size information in the data quality assessment matrix; multiplying the reliability scores by the normalized weights of the sample size to obtain dynamic weight allocation coefficients; weighting and merging toxicity detection data from different sources based on the dynamic weight allocation coefficients; calculating the weighted average effect size using the DerSimonian-Laird random effects model to obtain heterogeneity correction results; fusing the heterogeneity correction results with the microplastic physicochemical characteristic parameters; and classifying and organizing the data according to the toxicity endpoint type to obtain the quality-corrected toxicity data. Step S3: Determine whether the correlation between acute toxicity and chronic toxicity in the quality-corrected toxicity data meets the biological constraints. If it does, perform collaborative learning training based on the constraint relationship between multiple toxic endpoints to establish an association mapping between physicochemical features and multiple toxic endpoints, and obtain a constraint fusion prediction algorithm. This includes: calculating the Pearson correlation between the acute toxicity data and chronic toxicity data in the quality-corrected toxicity data, determining whether the correlation coefficient is greater than a preset biological constraint threshold, and confirming that the biological constraints are met if it is. Obtaining the toxic endpoint correlation verification result; establishing a multi-task learning loss function based on the toxic endpoint correlation verification result, and weighting and combining the loss terms of acute toxicity, chronic toxicity, bioaccumulation, and genotoxicity to obtain a constraint loss function; training a neural network on the physicochemical feature parameters of the microplastics and the toxic endpoint data based on the constraint loss function, using a shared hidden layer to extract general feature representations, and obtaining a multi-toxic endpoint mapping network; integrating and optimizing the multi-toxic endpoint mapping network with the biological constraints, adding regularization constraint terms between toxic endpoints, and obtaining the constraint fusion prediction algorithm. Step S4: Input the physicochemical characteristic parameters of the sample to be tested into the constrained fusion prediction algorithm to calculate toxicity, and determine whether the prediction uncertainty exceeds the set threshold. If it does not exceed the threshold, combine Bayesian quantization to output the toxicity prediction value and confidence interval. Step S5: Calculate the risk quotient and analyze the mechanism contribution based on the predicted toxicity value, determine whether the risk quotient exceeds the safety threshold, and if it does, mark it as high risk and generate early warning information to obtain the graded risk assessment result.
2. The microplastic toxicity prediction method integrating machine learning and meta-analysis according to claim 1, characterized in that, Step S1 includes: Fourier transform infrared spectroscopy was used to perform full-band spectral scanning on microplastic samples. Based on the characteristic peak intensity ratio, the matching degree with the standard polymer spectrum was calculated to obtain chemical composition detection data. The microplastic sample was examined by scanning electron microscopy. The particle outline was automatically identified and the geometric parameters were measured. The equivalent circle diameter and aspect ratio were calculated to obtain morphology detection data. The electrophoretic mobility of the microplastic sample under various pH conditions was measured based on dynamic light scattering, and the Zeta potential distribution was calculated according to the Henry's function formula to obtain surface charge detection data. The chemical composition detection data, morphology detection data, and surface charge detection data are integrated according to a standardized format to obtain the physicochemical characteristic parameters of the microplastic.
3. The microplastic toxicity prediction method integrating machine learning and meta-analysis according to claim 1, characterized in that, The process involves calculating the Pearson correlation between the acute toxicity data and chronic toxicity data in the quality-corrected toxicity data, determining whether the correlation coefficient is greater than a preset biological constraint threshold, and confirming that the biological constraint condition is met if it is. This yields the toxicity endpoint correlation verification result, including: Logarithmic transformation was performed on the acute toxicity LC50 value and chronic toxicity NOEC value in the quality-corrected toxicity data to convert the values into a normal distribution form, thus obtaining standardized toxicity data pairs. Based on the standardized toxicity data, the Pearson correlation coefficient was calculated, and the correlation was quantified by the formula of covariance divided by standard deviation to obtain the correlation coefficient of the toxicity endpoint. The correlation coefficient of the toxicity endpoint is numerically compared with the biological constraint threshold of 0.6 to determine whether the acute toxicity and chronic toxicity meet the positive correlation requirement, and the constraint condition determination result is obtained. Based on the result of the constraint judgment, determine whether to perform multi-toxicity endpoint collaborative learning. If the constraint is met, the data is marked as qualified and enters the subsequent training process to obtain the correlation verification result of the toxicity endpoint.
4. The microplastic toxicity prediction method integrating machine learning and meta-analysis according to claim 1, characterized in that, Step S4 includes: The physicochemical characteristic parameters of the sample to be tested are input into the constrained fusion prediction algorithm for forward propagation calculation. The predicted values of acute toxicity, chronic toxicity, bioaccumulation and genotoxicity are calculated through the shared hidden layer and the task-specific output layer, respectively, to obtain multi-endpoint toxicity prediction data. Based on Monte Carlo Dropout sampling, the multi-endpoint toxicity prediction data is forwarded multiple times to calculate the prediction variance and cognitive uncertainty, and the prediction uncertainty quantification result is obtained. The prediction uncertainty quantification result is compared with a preset uncertainty threshold to determine whether the total uncertainty exceeds the set threshold. If it does not exceed the threshold, the prediction result is confirmed to be reliable, and a reliability determination result is obtained. Based on the reliability determination results and the Bayesian posterior distribution calculation, the confidence interval of the multi-terminal toxicity prediction value is quantified to obtain the prediction result containing the toxicity prediction value and the confidence interval.
5. The microplastic toxicity prediction method integrating machine learning and meta-analysis according to claim 1, characterized in that, Step S5 includes: Based on the predicted toxicity value and the predicted no-effect concentration, the risk quotient is calculated. The predicted environmental concentration is divided by the predicted no-effect concentration to obtain the risk quotient data. The risk level is classified according to the size of the risk quotient to obtain the quantitative risk assessment result. The predicted toxicity values are analyzed and calculated using SHAP value analysis to quantify the contribution of each physicochemical characteristic parameter to the toxicity prediction results, identify the dominant toxicity mechanism and key influencing factors, and obtain the mechanism contribution analysis results. The risk quotient in the quantitative risk assessment result is compared with the safety threshold of 1.0 to determine whether the risk quotient exceeds the safety threshold. If it does, it is marked as a high-risk level and an early warning mechanism is triggered to obtain the risk level determination result. A standardized assessment report is generated based on the risk level determination results and mechanism contribution analysis results. The report integrates the risk level, toxicity mechanism explanation, and management recommendations to obtain the graded risk assessment results.
Citation Information
Patent Citations
Construction method of micro-plastic toxicity prediction model based on transcriptomics and QSAR (Quantitative Synthetic Aperture Radar) model
CN120015125A