Method for constructing a pfas risk prediction model based on kocs and bcf

CN121506267BActive Publication Date: 2026-08-18KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511619538.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-06
Publication Date
2026-08-18
Estimated Expiration
2045-11-06

AI Technical Summary

Technical Problem

值得关注的是,PFASs在食物链中表现出显著的生物放大效应,当其在生物体内蓄积浓度达到一定阈值时,可引发多种毒效应,对人类健康构成严重威胁

Benefits of technology

[0026] Compared with existing technologies, the beneficial effects of this invention are as follows: As emerging persistent organic pollutants, the prediction of the environmental migration and fate of PFASs is a key challenge in environmental risk assessment. Traditional ST-QSPR models cannot fully analyze the complex partitioning behavior of PFASs in multi-media environmental systems, making them unsuitable for practical risk assessment needs. This invention constructs an MT-MLR-QSPR model based on the molecular structure characteristics of 14 representative PFASs (covering 9 PFCAs, 4 PFSAs, and 1 FOSAs) that can simultaneously predict the organic carbon-water partition coefficient (log Koc) and bioaccumulation factor (log BCF). The comprehensive validation and evaluation results of the model show that its goodness of fit (R² of the PFASs-log Koc model) is high. 2 =0.873, R² of the PFASs-log BCF model2 =0.738), robustness (PFASs-log Koc model) =0.932, PFASs-log BCF model =0.702) and external predictive power (PFASs-log Koc model and PFAs-log BCF model) , , All values ​​were greater than 0.5, meeting the QSPR model application criteria. Model mechanism analysis revealed that the atomic distance (ADC) closest to the molecular geometric center, molecular polarity index (MPI), and nonpolar surface area (NPSA) of PFASs are the main structural factors influencing the partitioning and bioaccumulation behavior of PFASs in the solid phase (soil/sediment)-aquatic phase. Specifically, log Koc is mainly dominated by the antagonistic effect of MPI and NPSA; log BCF is influenced by the synergistic effect of ADC and MPI; and carbon chain length is a common key driving factor affecting the partitioning behavior of both. The results of this study are of great significance for a deeper understanding of the migration patterns and environmental enrichment characteristics of PFASs between sediment/soil and aquatic phases, providing crucial basic data support for subsequent environmental risk assessments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121506267B_ABST
    Figure CN121506267B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of pollutant risk prediction, and discloses a method for constructing a PFASs risk prediction model based on Koc and BCF, comprising the following steps: collecting various PFASs data, processing the data to serve as training data for modeling; calculating molecular descriptors of various PFASs, and preliminarily screening the molecular descriptors through a Pearson correlation analysis method, and further screening the preliminarily screened molecular descriptors through a multi-task elastic net regression; and using the screened molecular descriptors to construct a PFASs prediction model based on a multi-task multivariate linear regression algorithm combined with a multi-task joint forward stepwise selection strategy. The purpose of the present application is to construct a prediction model capable of synchronously predicting an organic carbon-water partition coefficient log Koc and a biological enrichment factor log BCF.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of pollutant risk prediction technology, and in particular to a method for constructing a PFASs risk prediction model based on Koc and BCF. Background Technology

[0002] Perfluorinated and polyfluoroalkyl substances (PFASs) are a typical class of emerging persistent organic pollutants (POPs), and their widespread environmental presence and potential health risks have become a research hotspot in the global environmental science field. PFASs are diverse, mainly including perfluoroalkyl carboxylic acids (PFCAs), perfluoroalkyl sulfonic acids (PFSAs), perfluorooctanesulfonamides (FOSAs), and perfluorinated telomeric alcohol compounds (FTOHs). Due to their unique hydrophobic and oleophobic properties, thermal stability, and chemical inertness, PFASs are widely used in industrial production and everyday consumer products. Large-scale global production and use have led to the widespread migration and diffusion of PFASs in the environment. Studies have shown that these compounds can be widely distributed in various environmental media through multiple pathways, including water transport, atmospheric transport and deposition, and soil infiltration, and can persistently accumulate in water sediments, soil particles, and organisms. Of particular concern is the significant biomagnification effect of PFASs in the food chain; when their concentration in organisms reaches a certain threshold, they can trigger various toxic effects, posing a serious threat to human health. Therefore, a thorough analysis of the cross-media migration behavior, bioaccumulation patterns, and transformation mechanisms of PFASs in environmental systems is of significant scientific importance for scientifically predicting their long-term environmental fate, constructing accurate ecological and health risk assessment systems, and formulating effective control strategies. Summary of the Invention

[0003] The purpose of this invention is to construct a predictive model that can simultaneously predict the organic carbon-water partition coefficient (log Koc) and the bioaccumulation factor (log BCF), and to provide a method for constructing a PFASs risk prediction model based on Koc and BCF.

[0004] To achieve the above-mentioned objectives, the embodiments of the present invention provide the following technical solutions:

[0005] The method for constructing PFASs risk prediction models based on Koc and BCF is characterized by the following steps:

[0006] Step 1: Collect various PFAS data, process them, and use them as training data for modeling.

[0007] Step 2: Calculate molecular descriptors for various PFASs, and perform preliminary screening of molecular descriptors using Pearson correlation analysis. Then, further screen the molecular descriptors after preliminary screening using multi-task elastic network regression.

[0008] Step 3: Using the selected molecular descriptors, a PFASs prediction model is constructed based on a multi-task multiple linear regression algorithm and a multi-task joint forward stepwise selection strategy.

[0009] Furthermore, in step 1, the collected PFASs include: PFBA, PFPeA, PFHxA, PFHpA, PFOA, PFNA, PFDA, PFUnDA, PFDoDA, PFBS, PFHxS, PFOS, PFDS, and PHOSA.

[0010] Furthermore, step 2, the step of calculating the molecular descriptors of various PFASs, includes:

[0011] The molecular structures of all PFASs were drawn using Chem Draw 22.0 software, and the molecular structures of all PFASs were pre-optimized using Chem 3D 22.0 software based on molecular theory under the MM2 force field.

[0012] The molecular structure of PFASs in the neutral electronic ground state was deeply optimized using the B3LYP / 6-31G* algorithm in the Gaussian 16 package. Frequency analysis confirmed that the obtained structure had no imaginary frequency, ensuring that the stable molecular configuration with the lowest energy was obtained.

[0013] The optimized molecular structure of PFASs was calculated using the Multiwfn program, and several molecular descriptors were obtained from various PFASs. These descriptors include the sphericity, density, area of ​​the positive potential region, variance of the positive potential region, distance of the nearest atom to the geometric center of the molecule, nonpolar surface area, and molecular polarity index of the PFASs.

[0014] Furthermore, step 2, the step of screening molecular descriptors using multi-task elastic network regression, includes:

[0015] A multi-task elastic network regression algorithm based on a regularization strategy is used to further filter multiple molecular descriptors selected by Pearson correlation analysis; the multi-task elastic network regression algorithm adjusts hyperparameters through Bayesian optimization. By combining L1 and L2 penalties to achieve dual constraints, the objective function of the multi-task elastic net regression algorithm is constructed as follows:

[0016] Where W=[w Koc / w BCF ] is the coefficient matrix; w Koc and w BCF These represent the weight vectors for the log Koc and log BCF tasks, respectively. For sparsity penalty, w j It is the weight of the j-th feature. The L1 penalty for the W matrix; For Frobenius norm, w jt It is the weight of the j-th feature in the t-th sample. This is the L2 penalty of the W matrix; p is the number of features, j=1,2,...,p; T is the number of samples, t=1,2,...,T; This is the regularization strength control coefficient; For hyperparameters; y t Let X be the true target value of the t-th sample; X is the feature matrix; w t Let be the weight vector of the t-th sample.

[0017] Furthermore, in step 3, the formula for the multi-task multiple linear regression algorithm is:

[0018] Among them, X i Let X be the i-th molecule descriptor; Score(X) i ) represents the score of the i-th molecular descriptor associated with log Koc or log BCF; , C is the scaling factor; orr ( ) represents the correlation coefficient; Var() represents the variance; ΔMSE Koc ΔMSE represents the change in the mean squared error of log Koc. BCF This represents the change in the mean square error of log BCF.

[0019] Furthermore, step 4 is included to validate the PFASs prediction model; using the coefficient of determination R... 2 The root mean square error (MSE) is used to evaluate the goodness of fit of the model.

[0020] Furthermore, it also includes step 5, which analyzes the application domain of the PFASs prediction model;

[0021] Williams plots are used to visualize the application domain of the PFASs prediction model, and the residuals are standardized. The joint analysis with the leverage value h gives statistical confidence to the applicability of the PFASs prediction model;

[0022] Williams plot derived from standardized residuals Composed of the leverage value h, Reflecting the degree of deviation between predicted and experimental values, when When this occurs, it indicates that the predicted value of the sample is an outlier. The calculation formula is:

[0023] in, and , respectively, are the experimental and predicted values ​​of the i-th PFAS compound; n is the number of PFAS compounds in the modeling dataset; k is the number of molecular descriptors of PFAS applicable in the PFAS prediction model;

[0024] h is used to measure the similarity between the molecular structure of PFASs and the training set. When h exceeds the warning leverage value h*, it indicates that the structure of the sample is significantly different from the structure of the samples in the training set. The formulas for calculating h and h* are as follows:

[0025] Where, x i Let be the vector of the i-th PFASs molecular descriptor; T is the matrix transpose; X is the matrix composed of PFASs molecular descriptors in the training set.

[0026] Compared with existing technologies, the beneficial effects of this invention are as follows: As emerging persistent organic pollutants, the prediction of the environmental migration and fate of PFASs is a key challenge in environmental risk assessment. Traditional ST-QSPR models cannot fully analyze the complex partitioning behavior of PFASs in multi-media environmental systems, making them unsuitable for practical risk assessment needs. This invention constructs an MT-MLR-QSPR model based on the molecular structure characteristics of 14 representative PFASs (covering 9 PFCAs, 4 PFSAs, and 1 FOSAs) that can simultaneously predict the organic carbon-water partition coefficient (log Koc) and bioaccumulation factor (log BCF). The comprehensive validation and evaluation results of the model show that its goodness of fit (R² of the PFASs-log Koc model) is high. 2 =0.873, R² of the PFASs-log BCF model2 =0.738), robustness (PFASs-log Koc model) =0.932, PFASs-log BCF model =0.702) and external predictive power (PFASs-log Koc model and PFAs-log BCF model) , , All values ​​were greater than 0.5, meeting the QSPR model application criteria. Model mechanism analysis revealed that the atomic distance (ADC) closest to the molecular geometric center, molecular polarity index (MPI), and nonpolar surface area (NPSA) of PFASs are the main structural factors influencing the partitioning and bioaccumulation behavior of PFASs in the solid phase (soil / sediment)-aquatic phase. Specifically, log Koc is mainly dominated by the antagonistic effect of MPI and NPSA; log BCF is influenced by the synergistic effect of ADC and MPI; and carbon chain length is a common key driving factor affecting the partitioning behavior of both. The results of this study are of great significance for a deeper understanding of the migration patterns and environmental enrichment characteristics of PFASs between sediment / soil and aquatic phases, providing crucial basic data support for subsequent environmental risk assessments. Attached Figure Description

[0027] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0028] Figure 1 This is a framework diagram of the method of the present invention;

[0029] Figure 2 This is the QQ image in Embodiment 1 of the present invention, wherein, Figure 2 (a) in the figure is the Q-Q graph of log Koc. Figure 2 (b) in the figure is the QQ graph of log BCF;

[0030] Figure 3(a) is a heat map of log Koc in Embodiment 1 of the present invention;

[0031] Figure 3(b) is a heat map of log BCF in Embodiment 1 of the present invention;

[0032] Figure 4 This is a graph showing the relationship between the performance of the PFASs prediction model and the number of molecular descriptors in Example 2 of the present invention.

[0033] Figure 5The correlation between the predicted and experimental values ​​of PFASs' log Koc and log BCF in Example 2 of this invention is shown. Figure 5 In the table, (a) represents the correlation between the log Koc predicted values ​​and the experimental values ​​of PFASs. Figure 5 (b) in the figure represents the correlation between the logBCF predicted values ​​and the experimental values ​​of PFASs;

[0034] Figure 6 The diagram shows the log Koc and log BCF residual distributions of the PFASs prediction model in Embodiment 2 of this invention. Figure 6 (a) in the figure is the log Koc residual distribution of the PFASs prediction model. Figure 6 (b) in the figure is the log BCF residual distribution of the PFASs prediction model;

[0035] Figure 7 The figures show Williams plots of log Koc and log BCF for the PFASs prediction model of this invention, where... Figure 7 (a) in the figure is the Williams plot of the log Koc of the PFAS prediction model. Figure 7 (b) in the figure is the Williams plot of the log BCF of the PFASs prediction model. Detailed Implementation

[0036] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0037] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this invention, the terms "first," "second," etc., are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance, or suggesting any such actual relationship or order between these entities or operations. Additionally, the terms "connected," "linked," etc., can refer to a direct connection between elements or an indirect connection via other elements.

[0038] Example 1:

[0039] Environmental media partition coefficients are important parameters describing the distribution behavior of chemical substances among different environmental media such as water, air, soil, and sediment. Among them, the organic carbon-water partition coefficient (Koc) is a key indicator, quantitatively characterizing the distribution balance of organic pollutants between the solid phase (soil / sediment) and the liquid phase (water body). Its value directly reflects the adsorption tendency of pollutants at the soil-water and sediment-water interfaces. The bioconcentration factor (BCF) describes the balance process of organic matter between the environmental media (mainly the aquatic phase) and the biological organic phase (mainly fish), directly reflecting the potential for the absorption and accumulation of organic compounds in water by organisms. It is a core indicator for quantifying the migration of pollutants along the "water-biological phase." Koc and BCF are important parameters for assessing the environmental fate and ecological risk of organic pollutants. Studying Koc and BCF of PFASs is crucial for accurately analyzing their migration and distribution behavior in environmental media (such as water, soil, and sediment) and their accumulation characteristics in organisms. They have a synergistic effect in the environmental and ecological risk assessment of PFASs from the two dimensions of "migration in environmental media" and "accumulation through ecological exposure", respectively.

[0040] However, current laboratory studies on the determination of Koc and BCF in PFASs still face many challenges. The primary constraint is the operational procedures, which typically rely on costly equipment, long experimental cycles, and complex processes, resulting in low efficiency and poor operability. Secondly, BCF determination involves animal testing, raising certain ecological and ethical concerns. To address these challenges, researchers are committed to exploring more efficient and economical model prediction methods. Early studies primarily constructed prediction models through three approaches: models based on PFAS carbon chain length, models based on PFAS physicochemical properties, and models based on fragment constants. However, these models have limitations in application, mainly in that they are usually applicable to a limited range of PFASs, and their external predictive capabilities are generally weak, making it difficult to meet the demand for accurate prediction of PFAS environmental behavior under complex and variable real-world environmental conditions.

[0041] With the in-depth development of computational chemistry and cheminformatics, quantitative structure-property relationship (QSPR) models have become a powerful and widely used theoretical prediction tool. QSPR (also known as quantitative structure-property relationship) establishes a mathematical functional relationship between a compound's molecular structure descriptor and its target physicochemical properties, enabling rapid and efficient prediction of properties. This model has significant advantages such as high stability, strong universality, and clear physicochemical meaning of the molecular descriptor, providing a solid theoretical foundation and reliable computational means for predicting the properties of complex compounds. Based on these characteristics and advantages, this method has been widely applied to the prediction of partition coefficients of PFASs in various environmental media. Researchers have successfully constructed various linear and nonlinear QSPR models by systematically screening suitable molecular descriptors as input variables. The results show that the determination coefficients (R²) of these models are... 2 The values ​​are mostly concentrated between 0.7 and 0.9, and the root mean square error (RMSE) is usually below 0.6, generally showing good fit and predictive ability.

[0042] Although significant progress has been made in recent years in QSPR model research on the partition coefficients of PFASs in environmental media, in terms of model stability and prediction accuracy, existing research still mainly focuses on single-task predictions. These studies typically focus on predicting a single partition coefficient (such as log Koc, log Koa, log Kow, etc.) using molecular descriptors. Single-task (ST) models have advantages in specific scenarios due to their simple structure, clear objectives, and clear relationships between independent and dependent variables. However, the partitioning behavior of PFASs in real-world environments is a comprehensive result of complex interactions between multiple media and interfaces. ST models have inherent limitations in resolving such complex correlations, namely, they struggle to simultaneously consider the potential connections and synergistic effects between multiple related properties. For example, a key molecular structure of a PFAS (such as carbon chain length, functional group type, etc.) not only affects its adsorption behavior on solid-phase organic carbon but is also closely related to its enrichment level in organisms. These properties do not exist in isolation but are governed by common structural factors. ST models typically only capture the influencing factors of a specific property, making it difficult to comprehensively reveal the systemic mechanism by which molecular structure affects the overall environmental behavior of PFASs. Given the limitations of ST models in revealing the complex environmental behavior correlations of PFASs, developing comprehensive models that can simultaneously consider the interaction of multiple factors and integrate and predict multiple related properties is crucial for accurately analyzing the actual migration and fate patterns of PFASs in the environment.

[0043] Multi-task learning (MTL), a machine learning paradigm capable of processing multiple related tasks in parallel, has been widely applied in fields such as computer vision, bioinformatics, health informatics, language, natural language processing, and networks. Its core advantage lies in enhancing overall model performance by coupling multiple tasks, creating shared information paths, and fully exploring the commonalities and correlations between tasks. Zhang et al. established ST and MTL models using Partial Least Squares Regression (PLSR), Random Forest (RF), and Deep Neural Network (DNN) algorithms, respectively. Their results showed that the MTL model generally outperformed the corresponding ST model. This is mainly due to MTL's ability to learn richer feature representations, effectively capture the commonalities of different tasks, and uncover potential patterns in the data, thereby enhancing its ability to analyze complex data. Furthermore, the flexibility and scalability of MTL also provide potential for applications in a wider range of fields. However, despite significant progress in other disciplines, the application of MTL in environmental science is still in its early stages, with relatively limited research.

[0044] PFASs exhibit highly complex migration and transformation behaviors in environmental systems, making accurate assessment of their environmental risks a significant challenge. To deeply analyze the environmental fate characteristics and migration patterns of PFASs, this proposal employs the MTL method to construct an MT-QSPR model capable of simultaneously predicting the Koc (Korea Common Occurrence) and BCF (Biochemical Flux) of PFASs. This model aims to fully utilize and explore the intrinsic correlation between Koc and BCF prediction tasks, identify key molecular descriptors that simultaneously influence Koc and BCF, improve the model's computational efficiency, and expand its application scenarios. This proposal provides crucial scientific evidence for quantifying the ecological and environmental risks of PFASs, supporting the formulation of more effective pollution control strategies. The MT-QSPR model framework proposed in this proposal can provide methodological references for other research in environmental science involving multi-objective and complex interrelationships (such as co-migration of multiple pollutants and simulation of multi-media environmental behavior), and by promoting the deep integration of machine learning and environmental science, it offers new ideas for solving complex environmental problems.

[0045] This invention is achieved through the following technical solutions, such as... Figure 1 As shown, a method for constructing a PFASs risk prediction model based on Koc and BCF is proposed, including the following steps:

[0046] Step 1: Collect various PFAS data, process them, and use them as training data for modeling.

[0047] To construct a reliable QSPR model, it is first necessary to collect high-quality experimental data relevant to the target task as the basis for modeling. The modeling data in this scheme mainly came from two sources: (1) batch experiments conducted on deionized water systems based on literature; (2) published research literature on the bioaccumulation effect of PFASs on the livers of silver carp, tilapia, and snakehead. The compounds finally included in the modeling analysis included 9 PFCAs, 4 PFSAs, and 1 FOSA, totaling 14 PFASs. For details, please refer to Table 1.

[0048] Table 1 Information on 14 types of PFASs

[0049] Given that some of the raw data came from different experimental conditions and environments, the collected raw data underwent the following processing to reduce the uncertainty of the analysis:

[0050] (1) Data preprocessing: For multiple log Koc or log BCF experimental values ​​of the same PFAS, outliers that significantly deviate from the overall dataset are screened to ensure that the coefficient of variation (CV) of the remaining samples is less than or equal to 15%. Subsequently, the average of the remaining data is taken as the modeling dataset for the same PFAS. The coefficients of variation of log Koc and log BCF data in the modeling dataset of the same PFAS are controlled below 15% to meet the QSPR model's requirement for data homogeneity.

[0051] (2) Preprocessed data validation: Statistical validation was performed on the preprocessed modeling dataset. The results showed that the log Koc values ​​of the 14 PFASs ranged from 1.31 to 4.15, with a range of 2.84 and an average of 2.62, with a corresponding standard deviation (SD) of 0.96; the log BCF values ​​of the 14 PFASs ranged from 2.03 to 5.08, with a range of 3.05 and an average of 3.80, with a corresponding SD of 1.10.

[0052] After the above processing, 80% of the data in the modeling dataset was randomly selected using Excel software as the training set (11 types of PFASs) to build the QSPR model; the remaining 20% ​​of the data was used as the test set (3 types of PFASs) to perform external validation of the QSPR model.

[0053] Step 2: Calculate molecular descriptors for various PFASs, perform preliminary screening of molecular descriptors using Pearson correlation analysis, and then further screen the preliminary molecular descriptors using multi-task elastic network regression.

[0054] The appropriate selection of molecular descriptors and physicochemical parameters is a core step in constructing reliable prediction models. Molecular descriptors are measures of the structural characteristics of compounds; specific computational methods typically convert molecular structures into numerical values ​​that reflect their structural information, facilitating computer processing and analysis. This approach primarily selects quantum chemical descriptors for the QSPR model. Quantum chemical descriptors, obtained through electron density calculations, offer the advantage of rapid acquisition and accurate reflection of subtle differences in molecular structure. Therefore, this approach chooses quantum chemical descriptors as the core parameters for constructing the QSPR model of PFASs, ensuring that the descriptor system possesses both computational efficiency and structural resolution capabilities.

[0055] Step 2 specifically includes the following steps:

[0056] Step 2-1: Calculate the molecular descriptors for various PFASs.

[0057] First, the molecular structures of 14 PFASs were drawn using ChemDraw 22.0 software. Then, the molecular structures of these 14 PFASs were pre-optimized using Chem3D 22.0 software under the MM2 force field based on molecular theory. Next, the B3LYP / 6-31G* algorithm from the Gaussian 16 package was used for deep optimization of the molecular structures of the PFASs in the neutral electronic ground state. Frequency analysis confirmed that the obtained structures had no imaginary frequencies, ensuring the acquisition of the lowest-energy stable molecular configuration. Finally, the optimized PFAS molecular structures were calculated using the Multiwfn program, resulting in 57 molecular descriptors from the 14 PFASs. These molecular descriptors cover key physicochemical information such as the molecular structural characteristics, orbital energy levels, electronegativity, atomic charge, and polarity of the PFASs.

[0058] Step 2-2: Calculate the linear correlation between molecular descriptors and log Koc and log BCF using Pearson correlation analysis to perform preliminary screening of molecular descriptors.

[0059] In QSPR modeling, the selection of molecular descriptors is crucial. While increasing the number of molecular descriptors can improve the model's goodness of fit, redundant descriptors can reduce the model's robustness and predictive ability. This approach uses Pearson correlation analysis to select molecular descriptors. Correlation analysis aims to assess the degree of linear association between molecular descriptors (autocorrelation) and between molecular descriptors and the target tasks (log Koc and log BCF). By calculating the linear autocorrelation between molecular descriptors and their linear association with the target tasks (log Koc and log BCF) using Pearson correlation analysis, key variables can be accurately identified. This method requires variables to follow a normal distribution. In this embodiment, QQ plots are used to verify that the data distribution conforms to the normal distribution characteristics. The QQ plot results for log Koc and log BCF are shown below. Figure 2 As shown in (a) and (b), the log Koc and log BCF values ​​of PFASs are found to be basically located on both sides of the distribution line, indicating that the data of the variables studied in this scheme are normally distributed, which meets the application conditions of Pearson analysis.

[0060] In Pearson correlation analysis, the p-value is usually used to determine the significance level between two variables. When p ≤ 0.05, it indicates that the two variables are significantly correlated, and the strength of the correlation can be inferred from the Pearson coefficient. The larger the Pearson coefficient, the stronger the correlation; the sign of the coefficient indicates a positive or negative correlation between the related variables. In the screening of molecular descriptors, if the Pearson coefficient of two molecular descriptors is greater than 0.9, they are considered to be highly collinear. In this case, only molecular descriptors that are more strongly correlated with the target task (such as log Koc, log BCF) and have a clear physical interpretation are retained to avoid multicollinearity. Based on the above principles, nine molecular descriptors that are significantly correlated with both log Koc and log BCF were finally identified. Their correlation heatmaps are shown in Figures 3(a) and 3(b), where red indicates a positive correlation and blue indicates a negative correlation; the darker the color, the stronger the correlation. Since some molecular descriptors are correlated with both log Koc and log BCF, a total of 14 molecular descriptors were screened after integration. For details, please refer to Table 2.

[0061] Table 2. 14 molecular descriptors of the screened PFASs

[0062] Steps 2-3 involve further filtering the initially selected molecular descriptors using a multi-task elastic network regression algorithm.

[0063] Given that the molecular descriptor data selected in this scheme is characterized by small sample size and high dimensionality, a multi-task elastic network regression (MT-EN) algorithm based on a regularization strategy is used to further refine the 14 initially selected molecular descriptors. The MT-EN algorithm adjusts hyperparameters through Bayesian optimization. It combines Lasso (L1 penalty) and Ridge (L2 penalty) to achieve dual constraints, and uses regularization terms to force different task models to share a similar structured parameter space. The objective function of MT-EN is constructed, and the selection results are used as input to the Multi-Task Multiple Linear Regression (MT-MLR) algorithm to support multi-task collaborative feature optimization.

[0064] The objective function of MT-EN is:

[0065] Where W=[w Koc / w BCF ] is the coefficient matrix; w Koc and w BCF These represent the weight vectors for the log Koc and log BCF tasks, respectively. For sparsity penalty, w j It is the weight of the j-th feature. The L1 penalty for the W matrix; For Frobenius norm, w jt It is the weight of the j-th feature in the t-th sample. This is the L2 penalty of the W matrix, used to handle collinearity; p is the number of features, j=1,2,...,p; T is the number of samples, t=1,2,...,T; This is the regularization strength control coefficient; y is a hyperparameter used to balance L1 / L2 weights, with a range of [0,1]; t Let be the true target value of the t-th sample; X is the feature matrix, where each row represents a sample and each column represents a feature; w t Let be the weight vector of the t-th sample, used to predict the target value.

[0066] After further screening of the 14 molecular descriptors using the MT-EN algorithm, a total of 7 molecular descriptors were finally obtained after integration.

[0067] Step 3: Using the selected molecular descriptors, a PFASs prediction model is constructed based on a multi-task multiple linear regression algorithm and a multi-task joint forward stepwise selection strategy.

[0068] This scheme uses the 7 molecular descriptors finally selected in step 2 as input to a multi-task multiple linear regression algorithm, calculates the score of each molecular descriptor, and performs linear regression on the molecular descriptors with scores greater than the threshold.

[0069] We employ a multi-task multiple linear regression (MT-MLR) algorithm combined with a multi-task forward selection strategy to optimize the parameters of the PFASs prediction model and construct the PFASs prediction model.

[0070] The following formula is for the multi-task multiple linear regression algorithm:

[0071] Among them, X i Let X be the i-th molecule descriptor; Score(X) i ) represents the score of the i-th molecular descriptor associated with log Koc or log BCF; , C is the scaling factor; orr ( ) represents the correlation coefficient, with a value range of [-1, 1]. A value of 1 indicates a perfect positive correlation, a value of -1 indicates a perfect negative correlation, and a value of 0 indicates no linear correlation; Var() represents the variance, used to measure the dispersion of the data distribution; ΔMSE Koc ΔMSE represents the change in the mean squared error of log Koc, used to evaluate the model's prediction error on the log Koc task. BCF This represents the change in the mean squared error of the log BCF, used to evaluate the model's prediction error on the log BCF task.

[0072] PFAS prediction models are based on independent modeling of shared features, establishing linear regression equations of the same form but with independent parameters for each prediction task:

[0073] Where Y1 is the prediction result of log Koc, and Y2 is the prediction result of log BCF; X1, X2, ..., X in Y1 n As the molecular descriptors most relevant to log Koc, X1, X2, ..., X in Y2 n This is the molecular descriptor most relevant to log BCF.

[0074] Since the MT-MLR algorithm selects 3 molecular descriptors each for log Koc and log BCF after score calculation, n=3 in Y1 and n=3 in Y2. It should be noted that the 3 molecular descriptors with the highest log Koc score and the 3 molecular descriptors with the highest log BCF score can be the same or different. That is, the 3 molecular descriptors with the highest log Koc score and the 3 molecular descriptors with the highest log BCF score can be the same. When the 3 molecular descriptors most important for predicting log Koc and log BCF are completely identical, this constitutes key evidence of employing a "multi-task joint forward stepwise selection strategy." This result clearly shows that feature selection is not performed independently for each task, but is based on a unified joint scoring function (which weightedly combines the improvement of features on the performance ΔMSE of the two task models), and through a forward stepwise iterative process, systematically selects a shared subset of features that is optimal for all tasks overall. Therefore, the core of this strategy—"joint" and "stepwise"—is fully realized under this condition: feature selection is performed jointly across tasks, ultimately producing a single shared feature set, and independent prediction models are built for each task based on this set.

[0075] Step 4: Validate the PFASs prediction model.

[0076] Goodness of fit reflects the model's ability to interpret training data. This approach uses the coefficient of determination (R²). 2 The goodness of fit of the model was evaluated using the root mean square error (MSE) and other metrics. To verify the robustness of the prediction model, leave-one-out cross-validation (LOOCV) was used for internal validation. This method is particularly suitable for small datasets, and the coefficient of determination in leave-one-out cross-validation was used. As a key indicator, A higher value indicates better model robustness. Meanwhile, the study utilizes independent test sets to externally validate the PFASs prediction model, using external validation metrics... Evaluate the predictive ability of the model.

[0077] Step 5: Analyze the application domain of the PFASs prediction model.

[0078] The reliability of the prediction results of the PFASs prediction model for the property data of compounds is directly related to the applicable scope of the PFASs prediction model. Usually, only the compounds within the application field of the PFASs prediction model can ensure the reliability of the prediction results of the PFASs prediction model for their property data. According to the OECD model construction criteria, in this solution, the Leverage method based on distance metric (i.e., Williams plot) is used to visually characterize the application domain of the PFASs prediction model. This method assigns statistical confidence to the applicable scope of the PFASs prediction model through the joint analysis of the standardized residuals and the leverage value (h).

[0079] The Williams plot consists of the standardized residuals and the leverage value (h), reflecting the deviation degree between the predicted value and the experimental value. When , it indicates that the predicted value of this sample is an outlier, and the calculation formula is:

[0080] where, and are the experimental value and the predicted value of the i-th PFASs compound respectively; n is the number of PFASs compounds in the modeling dataset; k is the number of molecular descriptors of PFASs applicable in the PFASs prediction model.

[0081] h is used to measure the similarity between the molecular structure of PFASs and the training set. When h exceeds the warning leverage value h*, it indicates that the structure of this sample is significantly different from the structure of the samples in the training set. The calculation formulas for h and h* are respectively:

[0082] where, x i is the vector of the i-th PFASs molecular descriptor; T is the matrix transpose; X is the matrix composed of PFASs molecular descriptors in the training set.

[0083] Based on this, the effective application domain of the PFASs prediction model is defined as the region that simultaneously satisfies and h < h*.

[0084] Example 2:

[0085] In this example, experimental verification is carried out on the basis of Example 1 above.

[0086] (1) Molecular descriptors screened based on the MT-EN algorithm.

[0087] Based on the Python 3.10 platform, this embodiment uses the MT-EN algorithm with 14 molecular descriptors obtained from the initial screening as independent variables and log Koc and log BCF as dependent variables for modeling. The model parameters are optimized using Bayesian methods. When the Lasso parameter combination is 0.0862 and 0.8893, the PFASs prediction model minimizes the negative mean squared error. The MT-EN algorithm, by simultaneously optimizing the loss functions of two objective tasks (prediction of log Koc and log BCF), selects 7 key molecular descriptors from 14 molecular descriptors: MPI, ICS, NPSA, ADC, ADF, qF−, and ω. cubic .

[0088] The performance of the PFAS prediction model was assessed by calculating the MSE and R² for each objective task (log Koc and log BCF) separately. 2 An evaluation was conducted, and the results are shown in Table 3.

[0089] Table 3 Statistical parameters of the MT-EN algorithm

[0090] (ii) PFAS prediction model based on MT-MLR algorithm.

[0091] Using Python 3.10 as the development platform, the eight molecular descriptors selected by the MT-EN algorithm were used as input to the MT-MLR algorithm. The MT-MLR algorithm used stepwise regression to determine the optimal combination of three molecular descriptors. To determine the optimal prediction model, this embodiment analyzed the performance metrics (R²) of the PFASs prediction model. 2 The trend of MSE (and MSE) with the increase of the number of selected molecular descriptors, such as Figure 4 As shown in the figure. The results indicate that when the number of molecular descriptors increases from 1 to 3, the MSE decreases significantly, reaching its lowest value (1.170) with 3 molecular descriptors. As the number of molecular descriptors continues to increase, the MSE increases slightly; R 2 The value increases as the number of molecular descriptors increases from 1 to 3, reaching its highest value (0.805) with 3 molecular descriptors, and then gradually decreases and stabilizes.

[0092] Based on the OECD model validation criteria, and considering both the goodness of fit and robustness of the model, the model containing three molecular descriptors was selected as the optimal MT-MLR-QSPR model for PFASs prediction, as shown below:

[0093] Where ADC is the atomic distance of PFASs to the nearest molecular geometric center; MPI is the molecular polarity index of PFASs; and NPSA is the nonpolar surface area of ​​PFASs.

[0094] (III) Evaluation and validation of the PFASs prediction model (i.e., the MT-MLR-QSPR model).

[0095] Tables 4 and 5 show the statistical parameters of the internal and external validations of the optimal MT-MLR-QSPR model, respectively.

[0096] Table 4 Internal validation statistics of the optimal MT-MLR-QSPR model

[0097] Table 5 External validation statistics of the optimal MT-MLR-QSPR model

[0098] Where, N train R is the number of PFAS species included in the training set. 2 The coefficient of determination; is the multiple correlation coefficient for leave-one-out cross-validation; MSE train is the mean squared error of the training set samples; , and This is the external validation metric; MSE text is the mean squared error of the test set samples. According to the evaluation criteria for the MT-MLR-QSPR model, when the model's statistical parameter R² > 0.6, , When the result is positive, it indicates that the model has good fit, robustness, and external predictive ability.

[0099] Analysis shows that all statistical parameters of the proposed PFASs prediction model meet the QSPR model evaluation criteria, demonstrating good goodness of fit, robustness, and external predictive ability, thus conforming to the requirements of the QSPR model construction criteria. Furthermore, the model's... This indicates that the model does not exhibit overfitting.

[0100] Table 6 presents the specific results of simultaneous prediction of the log Koc and log BCF values ​​of PFASs using the optimal MT-MLR-QSPR model.

[0101] Table 6. Numerical values ​​of molecular descriptors and prediction results in the optimal MT-MLR-QSPR model.

[0102] The correlations between the predicted and experimental values of log Koc and log BCF for PFASs are shown in Figure 5 (a) and (b) in it. It can be seen that all data points are distributed near the 45˚ line, indicating that the model has a high prediction accuracy for the log Koc and log BCF values of PFASs.

[0103] Figure 6 (a) and (b) in it show the residual distributions of the predicted log Koc and log BCF of the PFASs prediction model. It can be seen that all residuals are randomly distributed on both sides of the baseline without obvious regularity, indicating that there is no systematic error in the model.

[0104] (IV) Analysis of the application domain of the PFASs prediction model.

[0105] Based on the Williams plot method, the application domain of the optimal MT-MLR-QSPR model constructed in this embodiment was defined. According to the calculation of h and h*, the warning leverage value (h*) of the two-task model is 1.09 for both. As shown in Figure 7 (a) and (b) in it, all samples in the training set and the test set satisfy and the dual conditions of h < h*, and completely fall within the range of the model application domain. This result proves that the constructed MT-MLR-QSPR model has reliable prediction ability and good generalization ability within its application domain.

[0106] (V) Mechanism analysis of the PFASs prediction model.

[0107] In this embodiment, the MT-MLR-QSPR model was interpreted mechanistically to analyze the main structural factors affecting the log Koc and log BCF values of PFASs. The MT-MLR-QSPR model shows that the log Koc and log BCF values of PFASs have a certain correlation with ADC (the distance from the atom closest to the molecular geometric center), MPI (molecular polarity index), and NPSA (non-polar surface area). The standardized regression coefficients in the MT-MLR-QSPR model refer to the regression coefficients when all variables are represented in a standardized form. Because they use the same measurement unit, it makes the independent variables (molecular descriptors) more comparable.

[0108] After calculation, in the linear regression equation for predicting log Koc of the PFASs prediction model, the standardized regression coefficients of the molecular descriptors ADC, MPI, and NPSA are 0.0013, -1.0387, and 0.8998 respectively. By comparing the magnitudes of their absolute values, it can be seen that the influence of the three molecular descriptors on log Koc is MPI > NPSA > ADC.

[0109] MPI, as a quantitative indicator of molecular polarity, reflects the non-uniformity of charge distribution within molecules. The MT-MLR-QSPR model shows a significant negative correlation between the MPI of PFASs and the log Koc value, meaning that the higher the molecular polarity, the weaker the adsorption capacity of PFASs on soil / sediment. Molecular dynamics simulations indicate that highly polar PFAS molecules tend to interact with hydrogen bonds or ions in the aqueous phase, thereby reducing their binding energy with organic carbon in soil / sediment. Further analysis shows that polar head groups can inhibit the adsorption behavior of PFASs by reducing hydrophobicity and enhancing hydration. In this study, the MPI of nine PFCAs was significantly negatively correlated with the log Koc value. Furthermore, the polar head groups (sulfonic acid group and sulfonamide group) of four PFSAs and one FOSA also had a certain inhibitory effect on adsorption, with the order of inhibitory efficacy being: sulfonic acid group > sulfonamide group > carboxyl group. This difference is attributed to the fact that sulfonic acid and sulfonamide groups can more effectively reduce the hydrophobicity of molecules and enhance hydration, thus significantly inhibiting the adsorption of PFASs on soil / sediment.

[0110] NPSA is a key physicochemical parameter describing the hydrophobic properties of molecular surfaces, representing the total surface area of ​​the nonpolar portion of the molecule. PFASs consist of carbon chains and polar functional groups (such as carboxylic acid groups and sulfonic acid groups), with the carbon chains being the nonpolar portion. NPSA is mainly determined by the length and structure of the carbon chains. The longer the carbon chain, the greater the proportion of the nonpolar portion in the PFAS molecule, and the higher the NPSA. The MT-MLR-QSPR model shows a positive correlation between the NPSA of PFASs and their log Koc value, promoting the adsorption of PFASs on soil / sediments. In this scheme, the variation trend of NPSA and log Koc value of PFASs is consistent, that is, with the increase of the carbon chain length of PFASs, both NPSA and log Koc value show an increasing trend. When the number of carbon chains in PFASs increased from 4 to 10, the NPSA increased from 113.88 Ų to 290.84 Ų, and the log Koc increased from 1.88 to 4.13, showing a particularly pronounced increasing trend. This also explains the phenomenon in multiple studies where the log Koc value increases with the increase of PFAS carbon chain length, indicating that the distribution of PFASs between soil / sediment and water is influenced by hydrophobic interactions.

[0111] In the linear regression equation for predicting log BCF using the PFASs prediction model, the standardized regression coefficients of ADC, MPI, and NPSA are 0.6485, -0.3004, and 0.1072, respectively. By comparing their absolute values, it can be seen that the influence of the three molecular descriptors on logKoc is ADC > MPI > NPSA.

[0112] In molecular structure analysis, ADC (Adaptive Coefficient of Molecular Size) is used to quantify the distance of the nearest atom to the geometric center of a molecule. Studies have shown that the distance of the nearest atom from the geometric center, as a structural variable, is negatively correlated with molecular size. Larger PFASs exhibit stronger hydrophobicity, and hydrophobic interactions can drive adsorption between adsorbates and adsorbents. Because hydrophobic interactions drive adsorption behavior between adsorbates and adsorbents, larger and more hydrophobic PFASs increase the adsorbate energy required to form holes between water molecules, leading to a stronger hydrophobic repulsion force between water molecules on the PFAS molecule surface.

[0113] As mentioned above, the higher the MPI, the stronger the hydrophilicity of PFASs, and the easier they are to retain in the aqueous phase. The linear regression equation for predicting log Koc in the PFASs prediction model shows a negative correlation between MPI and the BCF value of PFASs, consistent with the above description. High MPI generally enhances the water solubility of PFASs, reduces their partitioning ability in lipid tissues, and accelerates biological metabolism and excretion, leading to a decrease in log BCF. In this scheme, short-chain PFASs (such as PFBA) generally have weaker bioaccumulation capacity than long-chain PFASs (such as PFOS) due to their higher polarity. Meanwhile, carbon chain length is a key factor affecting NPSA. Studies by Bhhatarai and Gramatica indicate that the log BCF of PFASs in rainbow trout is positively correlated with the carbon chain length of PFASs. This means that the stronger the hydrophobicity of PFASs, the easier they are to accumulate in rainbow trout tissues.

[0114] (vi) Analysis of the characteristics and advantages of PFASs prediction models.

[0115] This approach constructs an MT-MLR-QSPR model (i.e., a PFASs prediction model) based on the molecular structural features of 9 PFCAs, 4 PFSAs, and 1 FOSA, simultaneously predicting log Koc and log BCF. Compared to existing models, this approach utilizes hierarchical molecular descriptor screening via MT-EN and MT-MLR, multi-task collaborative modeling, and independent task-specific validation to fully explore the intrinsic correlation between the log Koc and log BCF prediction tasks. It delves into the cross-media migration behavior, bioaccumulation patterns, and transformation mechanisms of PFASs in environmental systems, significantly enhancing the model's ability to analyze complex environmental behaviors.

[0116] Currently, methods for selecting molecular descriptors mainly fall into three categories: dimensionality reduction methods, traditional methods based on multiple linear regression, and artificial intelligence methods based on search algorithms. These methods are widely used in descriptor selection for single-task QSPR modeling, but they have limitations when dealing with multi-task modeling needs. Dimensionality reduction methods compress the feature space through linear / nonlinear transformations, which reduces data dimensionality but destroys the physicochemical meaning of molecular descriptors; methods based on multiple linear regression select descriptors by establishing a linear mapping between the prediction target and feature variables, but they sever the potential correlations between different prediction tasks; intelligent search algorithms can optimize feature combinations, but they struggle to achieve collaborative optimization across multiple target tasks. These existing methods are essentially limited by the single-task optimization paradigm and cannot effectively capture the synergistic effects between multi-task prediction tasks. To overcome these limitations, this method uses the MT-EN algorithm, employing feature sharing, regularization constraints, and a hierarchical architecture to create a data foundation for the prediction model. MT-EN replaces the traditional independent screening strategy with a feature-sharing mechanism, preserving the cross-task physical meaning of molecular descriptors; it replaces fragmented parameter optimization by using regularization constraints to model the relevance of tasks; this method establishes a combined architecture to extract common and specific characteristics hierarchically, improving model interpretability while reducing computational complexity. Research shows that the feature dataset constructed by MT-EN effectively integrates key structural information such as the carbon chain structure, polar functional group distribution, and steric hindrance effects of PFASs, providing a high-quality data foundation for prediction models that combines physical meaning and statistical effectiveness.

[0117] This innovative approach constructs a hybrid MT-MLR-QSPR model prediction framework, successfully achieving collaborative modeling of log Koc and log BCF. This framework uses MT-EN to jointly screen key molecular descriptors, resolving the cross-task correlations of three core parameters: ADC, MPI, and NPSA, systematically revealing the cross-interfacial allocation mechanism of PFASs in solid-phase (soil / sediment)-aqueous-biological multi-media systems. In contrast, traditional ST models, due to their fragmentation of the intrinsic connections between multi-media behaviors, are insufficient for predicting complex environmental systems. Our results show that key structural parameters of PFASs, such as carbon chain length, have dual environmental effects: they influence bioaccumulation through NPSA (long-chain enrichment effect) and reduce solid-phase adsorption capacity by weakening MPI (long-chain adsorption attenuation effect). This antagonistic mechanism of environmental behavior driven by shared structural groups is difficult to capture simultaneously in traditional ST models. The core advantage of the MT-MLR-QSPR model lies in its multi-task collaborative learning mechanism. The model simultaneously analyzes the sharing patterns of log Koc and log BCF, identifies cross-media core descriptors such as ADC and MPI, and reveals the synergistic mechanism of solid-phase partitioning and bioaccumulation. Based on the physical meaning of molecular descriptors, it clarifies that structural features such as carbon chain length and polar functional groups govern the multi-media partitioning behavior of PFASs through hierarchical regulation of hydrophobicity and polarity. This predictive model framework establishes a unified explanatory path from molecular structure to multi-environmental media behavior, not only verifying the common molecular driving mechanism of PFASs environmental behavior, but also providing theoretical support for the dynamic simulation of pollutant cross-media migration. Meanwhile, the established MTL modeling framework shows significant potential for expanded applications. In the field of environmental science, it can be extended to predict the multi-media partitioning behavior of pollutants such as antibiotics, nanomaterials, and microplastics, supporting the analysis of interactions in complex pollution scenarios; constructing an environmental fate model covering the entire atmosphere-water-soil-biological chain for environmental risk assessment. In ecotoxicology research, it correlates and predicts multiple environmental exposures and toxicity endpoints of chemicals (such as aquatic toxicity, biodegradability, and bioaccumulation), constructing a more comprehensive ecotoxicological risk assessment model. In terms of mechanism analysis, this method can be applied to the study of pollutant transport mechanisms across biological barriers (such as intestinal absorption and the blood-brain barrier), establishing a quantitative correlation between physicochemical properties and transmembrane transport efficiency. This modeling approach demonstrates unique integration and analytical capabilities in handling the complexity and interrelationships of environmental systems, providing new ideas and methods for a deeper understanding of pollutant environmental behavior and its potential risks.

[0118] This approach verifies the model's performance and reliability through a task-based independent evaluation strategy. As shown in Tables 4 and 5, the PFASs-log Koc model (R... 2 =0.873) The goodness of fit is better than that of PFAss-log BCF (R 2=0.738) model. This performance difference is essentially an objective reflection of the differences in the environmental behavior mechanisms of pollutants and the data foundation. From the perspective of the mechanism of action, Koc, as a parameter characterizing the distribution tendency of pollutants between solid organic matter and water phase, is dominated by physicochemical distribution processes and is directly affected by molecular structure descriptors, and can be accurately quantified through linear free energy relationships. However, BCF reflects the ability of pollutants to accumulate from environmental media to organisms, involving complex physiological mechanisms such as active transport, metabolic transformation, and bile excretion. This nonlinear physiological process is strongly affected by species specificity and exposure conditions, and is difficult to fully characterize through simple molecular descriptors. At the data foundation level, Koc has mature standardized testing methods, low data dispersion, and a relatively complete experimental database, providing a high-quality data foundation for model training. However, BCF, due to species specificity and sensitivity to exposure conditions, results in large experimental data variability and data fragmentation, which limits the representativeness of the data and restricts the generalization ability of the model. Related studies have shown that the log Koc of short-chain PFASs can be stably predicted, but its log BCF varies significantly among different organisms. The results of this study demonstrate that molecular structure-based predictions of environmental allocation behavior (such as Koc) have higher accuracy and reliability than predictions of bioaccumulation behavior (such as BCF). This finding fundamentally explains the current state of PFAS risk systems: Koc-based environmental migration models far surpass BCF-based bioaccumulation models in both prediction accuracy and application maturity, while the latter still needs to overcome the quantitative bottlenecks brought about by the complexity of biological systems.

[0119] (vii) Conclusion.

[0120] As emerging persistent organic pollutants, the prediction of PFASs' environmental migration and fate is a key challenge in environmental risk assessment. Traditional ST-QSPR models are insufficient to fully analyze the complex partitioning behavior of PFASs in multi-media environmental systems, making them unsuitable for practical risk assessment needs. This scheme constructs an MT-MLR-QSPR model based on the molecular structure characteristics of 14 representative PFASs (covering 9 PFCAs, 4 PFSAs, and 1 FOSAs) that can simultaneously predict the organic carbon-water partition coefficient (log Koc) and bioaccumulation factor (log BCF). Comprehensive validation and evaluation results show that the model's goodness of fit (R² of the PFASs-logKoc model) is high. 2 =0.873, R² of the PFASs-log BCF model 2 =0.738), robustness (PFASs-log Koc model) PFASs-log BCF model ) and external predictive power (PFASs-log Koc model and PFASs-log BCF model) All values ​​were greater than 0.5, meeting the QSPR model application criteria. Model mechanism analysis revealed that the atomic distance (ADC) closest to the molecular geometric center, molecular polarity index (MPI), and nonpolar surface area (NPSA) of PFASs are the main structural factors influencing the partitioning and bioaccumulation behavior of PFASs in the solid phase (soil / sediment)-aquatic phase. Specifically, log Koc is mainly dominated by the antagonistic effect of MPI and NPSA; log BCF is influenced by the synergistic effect of ADC and MPI; and carbon chain length is a common key driving factor affecting the partitioning behavior of both. The results of this study are of great significance for a deeper understanding of the migration patterns and environmental enrichment characteristics of PFASs between sediment / soil and aquatic phases, providing crucial basic data support for subsequent environmental risk assessments.

[0121] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for constructing a PFASs risk prediction model based on Koc and BCF, characterized in that, Includes the following steps: Step 1: Collect various PFAS data, process them, and use them as training data for modeling. Step 2: Calculate molecular descriptors for various PFASs, and perform preliminary screening of molecular descriptors using Pearson correlation analysis. Then, further screen the molecular descriptors after preliminary screening using multi-task elastic network regression. Step 3: Using the selected molecular descriptors, a PFASs prediction model is constructed based on a multi-task multiple linear regression algorithm and a multi-task joint forward stepwise selection strategy. Step 3 specifically includes: The molecular descriptors finally selected in step 2 are used as input to the multi-task multiple linear regression algorithm. The score of each molecular descriptor is calculated, and linear regression is performed on the molecular descriptors with scores greater than the threshold. When the molecular descriptors most important for the predictions of log Koc and log BCF are exactly the same, it constitutes key evidence for employing a multi-task joint forward stepwise selection strategy. Feature selection is performed jointly across tasks, ultimately producing a single shared feature set, and based on this, independent PFASs prediction models are built for each task. PFAS prediction models are based on independent modeling of shared features, establishing linear regression equations with the same form but independent parameters for each prediction task.

2. The method for constructing a PFASs risk prediction model based on Koc and BCF according to claim 1, characterized in that, In step 1, the collected PFASs include: PFBA, PFPeA, PFHxA, PFHpA, PFOA, PFNA, PFDA, PFUnDA, PFDoDA, PFBS, PFHxS, PFOS, PFDS, and PHOSA.

3. The method for constructing a PFASs risk prediction model based on Koc and BCF according to claim 1, characterized in that, Step 2, which involves calculating the molecular descriptors of various PFASs, includes: The molecular structures of all PFASs were drawn using Chem Draw 22.0 software, and the molecular structures of all PFASs were pre-optimized using Chem 3D 22.0 software based on molecular theory under the MM2 force field. The molecular structure of PFASs in the neutral electronic ground state was deeply optimized using the B3LYP / 6-31G* algorithm in the Gaussian 16 package. Frequency analysis confirmed that the obtained structure had no imaginary frequency, ensuring that the stable molecular configuration with the lowest energy was obtained. The optimized molecular structure of PFASs was calculated using the Multiwfn program, and several molecular descriptors were obtained from various PFASs. These descriptors include the sphericity, density, area of ​​the positive potential region, variance of the positive potential region, distance of the nearest atom to the geometric center of the molecule, nonpolar surface area, and molecular polarity index of the PFASs.

4. The method for constructing a PFASs risk prediction model based on Koc and BCF according to claim 1, characterized in that, Step 2, the step of screening molecular descriptors using multi-task elastic network regression, includes: A multi-task elastic network regression algorithm based on a regularization strategy is used to further filter multiple molecular descriptors selected by Pearson correlation analysis; the multi-task elastic network regression algorithm adjusts hyperparameters through Bayesian optimization. By combining L1 and L2 penalties to achieve dual constraints, the objective function of the multi-task elastic net regression algorithm is constructed as follows: Where W=[w Koc / w BCF ] is the coefficient matrix; w Koc and w BCF These represent the weight vectors for the log Koc and log BCF tasks, respectively. For sparsity penalty, w j It is the weight of the j-th feature. The L1 penalty for the W matrix; For Frobenius norm, w jt It is the weight of the j-th feature in the t-th sample. This is the L2 penalty of the W matrix; p is the number of features, j=1,2,...,p; T is the number of samples, t=1,2,...,T; This is the regularization strength control coefficient; For hyperparameters; y t Let X be the true target value of the t-th sample; X is the feature matrix; w t Let be the weight vector of the t-th sample.

5. The method for constructing a PFASs risk prediction model based on Koc and BCF according to claim 1, characterized in that, In step 3, the formula for the multi-task multiple linear regression algorithm is: Among them, X i Let X be the i-th molecule descriptor; Score(X) i ) represents the score of the i-th molecular descriptor associated with log Koc or log BCF; , C is the scaling factor; orr ( ) represents the correlation coefficient; Var() represents the variance; ΔMSE Koc ΔMSE represents the change in the mean squared error of log Koc. BCF This represents the change in the mean square error of log BCF.

6. The method for constructing a PFASs risk prediction model based on Koc and BCF according to claim 1, characterized in that, The process also includes step 4, validating the PFASs prediction model; using the coefficient of determination R... 2 The root mean square error (MSE) is used to evaluate the goodness of fit of the model.

7. The method for constructing a PFASs risk prediction model based on Koc and BCF according to claim 1, characterized in that, It also includes step 5, which analyzes the application domain of the PFASs prediction model; Williams plots are used to visualize the application domain of the PFASs prediction model, and the residuals are standardized. The joint analysis with the leverage value h gives statistical confidence to the applicability of the PFASs prediction model; Williams plot derived from standardized residuals Composed of the leverage value h, Reflecting the degree of deviation between predicted and experimental values, when When this occurs, it indicates that the predicted value of the sample is an outlier. The calculation formula is: in, and , respectively, are the experimental and predicted values ​​of the i-th PFAS compound; n is the number of PFAS compounds in the modeling dataset; k is the number of molecular descriptors of PFAS applicable in the PFAS prediction model; h is used to measure the similarity between the molecular structure of PFASs and the training set. When h exceeds the warning leverage value h*, it indicates that the structure of the sample is significantly different from the structure of the samples in the training set. The formulas for calculating h and h* are as follows: Where, x i Let be the vector of the i-th PFASs molecular descriptor; T is the matrix transpose; X is the matrix composed of PFASs molecular descriptors in the training set.