A method and system for predicting the ability of nanofiltration membranes to remove organic micropollutants
By constructing a machine learning-based prediction model for nanofiltration membrane retention performance and using the retention performance of standardized solutes to characterize the separation characteristics of nanofiltration membranes, the problem of difficulty in quickly predicting the retention performance of nanofiltration membranes for organic micropollutants in existing technologies is solved, achieving a simple and efficient prediction effect.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WUHAN UNIV
- Filing Date
- 2026-07-02
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies struggle to quickly and effectively predict the retention performance of nanofiltration membranes for organic micropollutants, and existing methods require complex membrane parameter acquisition and experimental testing, making it difficult to meet the needs for rapid evaluation and efficient application of nanofiltration technology.
By constructing a machine learning-based nanofiltration membrane retention performance prediction model, the separation characteristics of the nanofiltration membrane are characterized by the retention performance of standardized solutes, and a performance-performance mapping model is established to achieve rapid prediction of the retention performance of organic micropollutants.
It enables rapid, simple, and efficient prediction of the retention performance of organic micropollutants without the need for additional membrane property parameters and experimental testing. It is highly adaptable and can provide reliable technical support in complex water treatment scenarios.
Smart Images

Figure CN122490002A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of environmental protection and water treatment technology, specifically to a method and system for predicting the ability of nanofiltration membranes to remove organic micropollutants. Background Technology
[0002] Nanofiltration membranes are a type of pressure-driven semi-permeable membrane with nanoscale pore size, and their separation performance lies between that of reverse osmosis membranes and ultrafiltration membranes. Due to their high selective rejection capacity for divalent ions and organic micropollutants, while also possessing advantages such as low operating pressure and high water flux, nanofiltration technology has been widely used in fields such as advanced drinking water treatment, industrial wastewater reuse, and environmental pollution control.
[0003] The retention performance of nanofiltration membranes for trace organic pollutants is a key parameter in nanofiltration membrane screening, process optimization, and engineering design. Currently, this performance is mainly obtained through experimental determination. However, organic micropollutants in the aquatic environment are diverse, with significant differences in molecular structure and physicochemical properties; at the same time, there are many commercially available nanofiltration membrane models, and different membrane materials and preparation processes lead to significant differences in their separation characteristics. Therefore, conducting experimental tests on different membrane-pollutant combinations often requires a significant amount of time, manpower, and resources, making it difficult to meet the needs of rapid evaluation and efficient application of nanofiltration technology.
[0004] In recent years, machine learning methods have been increasingly applied to the field of membrane separation performance prediction due to their advantages in modeling complex relationships. However, most existing studies require the introduction of membrane characteristic parameters such as pore size, charge, and hydrophilicity / hydrophobicity as model inputs. These parameters are not only difficult to obtain and involve complex testing processes, but the test results are also easily affected by experimental conditions, which limits the application of the models.
[0005] To address the aforementioned problems, this invention proposes a machine learning-based method for predicting the retention performance of organic micropollutants in nanofiltration membranes. This method, considering practical application scenarios of nanofiltration membranes, selects several representative selective solutes. The retention behavior of these solutes characterizes the overall separation properties of the nanofiltration membrane. Combined with the physicochemical properties of the target organic micropollutants, the model is trained using existing experimental data, enabling rapid prediction of the retention performance of organic micropollutants without the need for additional membrane property parameters or experimental testing. Compared to existing methods, the method provided by this invention avoids the complex process of acquiring and characterizing membrane parameters, offering advantages such as ease of operation, high prediction efficiency, and strong applicability. It provides effective technical support for nanofiltration membrane screening, process design, and practical engineering applications. Summary of the Invention
[0006] To address the shortcomings of existing technologies, this invention provides a method for predicting the ability of nanofiltration membranes to remove organic micropollutants based on standardized solute retention performance. This invention establishes a machine learning-based predictive model of the retention performance of nanofiltration membranes for organic micropollutants. The retention performance of nanofiltration membranes for standardized solutes is used to characterize the separation properties of the membrane, and this characteristic is used as a proxy variable to achieve generalized prediction of the retention performance of different organic micropollutants. The predictive model constructed in this invention exhibits good adaptability and robustness, maintaining high prediction accuracy even under unexpected events such as membrane performance degradation and changes in operating conditions, providing reliable technical support for the application of nanofiltration membranes in complex water treatment scenarios.
[0007] To achieve the above objectives, the specific technical solution of the present invention is as follows:
[0008] In a first aspect, the present invention provides a method for predicting the retention performance of nanofiltration membranes for trace organic matter, comprising the following steps:
[0009] Dataset construction: Collect data on the retention of standardized solutes by nanofiltration membranes, data on the retention of organic micropollutants, and the property characteristics of the organic micropollutants, and construct a dataset;
[0010] Model establishment: The machine learning model was trained using the dataset to establish a mapping model that predicts the retention performance of organic micropollutants based on the standardized solute retention performance;
[0011] Performance prediction: The retention data of the nanofiltration membrane under test for the standardized solute and the property characteristics data of the target organic micropollutant are input into the mapping model to predict the retention performance of the nanofiltration membrane under test for the target organic micropollutant.
[0012] Furthermore, the standardized solute includes at least one monovalent salt and at least one polyvalent salt. When the target organic micropollutant is a charged organic compound, the prediction model obtained by training the nanofiltration membrane for at least one monovalent salt and at least one polyvalent salt, as described above, exhibits good predictive performance.
[0013] Furthermore, the monovalent salts include: sodium chloride (NaCl) and potassium chloride (KCl); and / or, the polyvalent salts include: sodium sulfate (Na2SO4, magnesium sulfate (MgSO4)), calcium chloride (CaCl2), and magnesium chloride (MgCl2).
[0014] Furthermore, the standardized solute also includes: at least one small organic molecule with a certain degree of hydrophobicity and / or at least one neutral small molecule.
[0015] Furthermore, the organic small molecules with a certain degree of hydrophobicity refer to organic small molecules with log D greater than 3, including: carbamazepine, bisphenol A, ibuprofen, diclofenac, atrazine; and / or, the neutral small molecules include: glucose, polyethylene glycol, sucrose.
[0016] This invention constructs a diverse set of retention characteristics by selecting standardized solutes that represent different separation mechanisms, thereby achieving a comprehensive characterization of membrane separation behavior. Specifically, monovalent salts (such as NaCl) characterize size sieving effects; polyvalent salts (such as Na₂SO₄ or CaCl₂) reflect electrostatic interactions at the membrane interface; and hydrophobic organic small molecules and / or neutral small molecules are further introduced to improve the model's predictive performance. Specifically, hydrophobic organic small molecules (such as carbamazepine or bisphenol A) characterize hydrophobic interactions, while neutral small molecules (such as glucose) supplement information on mechanisms such as size repulsion and polar interactions. Through this diverse set of standardized solutes, this invention enables the experimental construction of a retention characteristic system that reflects the comprehensive separation characteristics of the membrane, thus providing a more comprehensive information basis for predicting the retention performance of nanofiltration membranes for trace organic pollutants.
[0017] This invention selects appropriate standardized solutes based on the retention mechanism of nanofiltration membranes; for example, for positively charged nanofiltration membranes to be tested, it is necessary to demonstrate the repulsive effect of positive charges, since divalent cations (such as Ca) 2+ Mg 2+ Since ions carry a higher charge, polyvalent salts containing divalent cations (such as CaCl2) are more suitable than polyvalent salts containing monovalent cations (such as Na2SO4). Conversely, for negatively charged nanofiltration membranes, ions are mainly retained by trapping negative charges, as divalent anions (such as SO42-) are less effective. 2- The charge of the polyvalent anion is strong, so polyvalent salts containing divalent anions (such as Na2SO4) react better to this effect than polyvalent salts containing monovalent anions (such as CaCl2).
[0018] For strongly charged target organic micropollutants, due to their strong charge, the main sieving mechanism is charge repulsion and pore size sieving. Therefore, collecting the retention data of a monovalent salt (such as NaCl) and a polyvalent salt (such as Na2SO4) can effectively reflect these two mechanisms of charge repulsion and pore size sieving.
[0019] In addition to charge repulsion and pore size sieving, hydrophobic interactions are also important in hydrophobic target organic micropollutants. Therefore, it is necessary to introduce the retention data of nanofiltration membranes for organic small molecules with certain hydrophobicity to reflect the hydrophilicity and hydrophobicity of nanofiltration membranes, thereby improving the predictive ability of the prediction model to predict the retention performance of other hydrophobic substances.
[0020] Furthermore, for a positively charged nanofiltration membrane to be tested, the standardized solute includes: at least one monovalent salt and at least one polyvalent salt containing a divalent cation; and / or, for a negatively charged nanofiltration membrane to be tested, the standardized solute includes: at least one monovalent salt and at least one polyvalent salt containing a divalent anion; and / or, for a charged target organic micropollutant, the standardized solute includes: one monovalent salt and one polyvalent salt; for a hydrophobic target organic micropollutant, the standardized solute includes: at least one monovalent salt, at least one polyvalent salt, and at least one small organic molecule with a certain degree of hydrophobicity.
[0021] Furthermore, the water contact angle of the membrane is used instead of the data on the rejection of the nanofiltration membrane for the organic small molecules with certain hydrophobicity.
[0022] Furthermore, the property characteristics data of the organic micropollutants include: molecular weight and the charge state (charge) and logarithmic partition coefficient (log D) at pH = 7. The charge state and logarithmic partition coefficient at pH = 7 are calculated using the Chemicalize platform. The property characteristics data selected in this invention—molecular weight, charge state, and logarithmic partition coefficient—represent the size, electrical properties, and hydrophilicity / hydrophobicity of the organic micropollutants, respectively, corresponding to the mechanisms of size sieving, charge repulsion, and hydrophobic interaction in nanofiltration membrane retention mechanisms.
[0023] Furthermore, the data set is constructed as follows: data on the retention of standardized solutes and organic micropollutants by nanofiltration membranes are obtained through literature review or experimental measurement; property characteristic data of organic micropollutants are obtained through public chemical databases; and the above data are preprocessed to construct the data set.
[0024] Furthermore, the preprocessing includes deduplication, mean imputation, or missing value processing.
[0025] Furthermore, the input features of the machine learning model include: nanofiltration membrane retention data of standardized solutes; property characteristic data of organic micropollutants; and the output of the machine learning model is nanofiltration membrane retention data of organic micropollutants.
[0026] Furthermore, the machine learning model is an ensemble learning model based on decision trees. The training process includes dataset partitioning, model training, and performance evaluation (using R2, MAE, and RMSE to evaluate the model), and the model's prediction accuracy is improved through hyperparameter optimization. In one example of this invention, the XGboost model is trained using the dataset to establish a mapping model that predicts the retention performance of organic micropollutants based on standardized solute retention performance.
[0027] Furthermore, the prediction step specifically includes: testing the retention data of the nanofiltration membrane to be tested for standardized solutes; inputting the retention data of the nanofiltration membrane to be tested for standardized solutes and the property characteristics data of the target pollutant into the mapping model; and outputting the prediction results of the retention performance of the target organic micropollutant.
[0028] In a second aspect, the present invention provides a system for predicting the retention performance of nanofiltration membranes for trace organic matter, comprising:
[0029] Data acquisition module: used to acquire data on the retention of standardized solutes and organic micropollutants by nanofiltration membranes, as well as the property characteristics data of organic micropollutants, and to construct a dataset;
[0030] Model building module: Used to train machine learning models using datasets and build a mapping model that predicts the retention performance of organic micropollutants based on standardized solute retention performance;
[0031] Retention performance prediction module: Used to predict the retention performance of the nanofiltration membrane under test for target organic micropollutants using a mapping model.
[0032] Thirdly, the present invention provides a device for predicting the retention performance of nanofiltration membranes for trace organic matter, comprising: one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method for predicting the retention performance of nanofiltration membranes for trace organic matter.
[0033] The core idea of this invention is that the transport and retention behavior of typical solutes in nanofiltration membranes is essentially controlled by multiple factors, including membrane pore structure, surface charge, and membrane-solute interactions. Therefore, the retention data of a rationally selected target solute can, to a certain extent, comprehensively characterize the separation characteristics of the nanofiltration membrane. Based on this, this invention, while retaining the physicochemical properties (molecular weight, charge state, hydrophobicity) of organic micropollutants, uses the retention performance of the nanofiltration membrane for standardized solutes to characterize the separation characteristics of the nanofiltration membrane, establishing a performance-to-performance machine learning prediction model. The key to the performance-to-performance machine learning prediction model framework lies in the selection of standardized solutes. An ideal target solute should simultaneously meet the following conditions: first, it should be readily available in the literature, supporting data integration across studies; second, the testing system should be relatively clear, with relatively unified experimental definitions, facilitating comparisons between results from different literatures; and third, it should be able to cover the main mechanisms of action in nanofiltration membrane separation, especially reflecting differences in membrane pore structure, membrane surface charge, and hydrophilicity / hydrophobicity from different perspectives. Compared to mixed-salt systems, single-salt systems have simpler ionic compositions, making it easier to correlate retention results with the separation behavior of a specific salt, and their physical meaning is clearer. In contrast, mixed-salt systems often involve competitive transport, charge shielding, and coupling between different ions, making it difficult to directly separate experimental results and attribute them to a single salt. Therefore, this invention selects commonly found single-salt retention data from the literature as membrane separation performance characteristics to improve the feasibility of data integration and comparative analysis between different studies. Furthermore, to better reveal pore size sieving and hydrophobic interactions, retention data for neutral and hydrophobic small-molecule organic molecules are added as supplementary data. The significance of this performance-performance framework lies not only in replacing explicit membrane parameters with performance characteristics but also in providing new ideas for the design and screening of alternative solutes. That is, the model input is not limited to traditional membrane characterization parameters but can be further expanded to a set of alternative characteristics that can cover the main separation mechanisms. The resulting characteristic system retains strong mechanistic relevance and is more conducive to rapid prediction and method promotion for practical separation tasks.
[0034] Compared with the prior art, the advantages of the present invention are:
[0035] (1) The present invention establishes a standardized alternative test set, so that the required nanofiltration membrane property characteristics data are only the easily obtainable rejection rate of single inorganic salts and simple solutes, and the water contact angle, without directly relying on the difficult-to-determine physicochemical properties of nanofiltration membranes.
[0036] (2) This invention uses machine learning methods to predict the retention effect of nanofiltration membranes on complex pollutants by utilizing the retention characteristics of nanofiltration membranes on simple solutes, thus avoiding the need for tedious laboratory tests.
[0037] (3) This invention establishes a predictive model for the retention of pollutants by nanofiltration membranes, and through data accumulation and multi-scenario generalization, helps researchers to predict and optimize the performance of nanofiltration membranes in different scenarios. Attached Figure Description
[0038] Figure 1 This is a schematic diagram of the method for predicting the ability of nanofiltration membranes to remove organic micropollutants based on standardized solute retention performance according to the present invention.
[0039] Figure 2 This is a schematic diagram of the machine learning training process in Example 1;
[0040] Figure 3 The above are the prediction results of the XGBoost model on the training and test sets in Example 1.
[0041] Figure 4 This is a graph showing the grouping and retention performance of organic micropollutants in the test set of Example 1;
[0042] Figure 5 This is a distribution chart of SHAP importance in Example 1;
[0043] Figure 6 This is a Spearman correlation matrix heatmap showing the relationship between the physicochemical properties of the nanofiltration membrane and the single salt rejection rate in Example 1.
[0044] Figure 7 The RMSE results of the models based on different input features in Example 1 on the test set;
[0045] Figure 8 This is the prediction result of the model on the test set after introducing WCA as an input feature based on 5 single salts in Example 1. Detailed Implementation
[0046] To enable those skilled in the art to clearly and completely understand the technical solution of the present invention, the present invention will be further described in detail below with reference to embodiments. Obviously, the embodiments described herein are only for explaining the present invention and are not intended to limit the scope of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0047] This invention provides a method for predicting the retention performance of nanofiltration membranes for trace organic matter, such as... Figure 1 As shown, it includes the following steps:
[0048] Dataset construction: Collect data on the retention of standardized solutes by nanofiltration membranes, data on the retention of organic micropollutants, and the property characteristics of the organic micropollutants to construct a dataset;
[0049] Model establishment: The machine learning model was trained using the dataset to establish a mapping model that predicts the retention performance of organic micropollutants based on the standardized solute retention performance;
[0050] Performance prediction: The retention data of the nanofiltration membrane under test for the standardized solute and the property characteristics data of the target organic micropollutant are input into the mapping model to predict the retention performance of the nanofiltration membrane under test for the target organic micropollutant.
[0051] In one embodiment, the standardized solute comprises: at least one monovalent salt and at least one polyvalent salt; wherein the monovalent salt comprises: sodium chloride (NaCl) and potassium chloride (KCl); and / or, the polyvalent salt comprises: sodium sulfate (Na2SO4), magnesium sulfate (MgSO4), calcium chloride (CaCl2), and magnesium chloride (MgCl2).
[0052] In one embodiment, the standardized solute further includes: at least one small organic molecule with a certain degree of hydrophobicity and / or at least one neutral small molecule; wherein, the small organic molecule with a certain degree of hydrophobicity refers to an organic small molecule with log D greater than 3, including: carbamazepine, bisphenol A, ibuprofen, diclofenac, atrazine; and / or, the neutral small molecule includes: glucose, polyethylene glycol, sucrose.
[0053] In one embodiment, for a positively charged nanofiltration membrane, the standardized solute comprises: at least one monovalent salt and at least one polyvalent salt containing a divalent cation; and / or, for a negatively charged nanofiltration membrane, the standardized solute comprises: at least one monovalent salt and at least one polyvalent salt containing a divalent anion; and / or, for a charged target organic micropollutant, the standardized solute comprises: a monovalent salt and a polyvalent salt; for a hydrophobic target organic micropollutant, the standardized solute comprises: at least one monovalent salt, at least one polyvalent salt, and at least one small organic molecule with a certain degree of hydrophobicity.
[0054] In one embodiment, the data on the rejection of the nanofiltration membrane for the organic small molecules with a certain degree of hydrophobicity is replaced by the water contact angle of the membrane.
[0055] In one embodiment, the property characteristics data of the organic micropollutant include: molecular weight (MW) and charge state (Charge) and logarithmic partition coefficient (log D) at pH=7.
[0056] In one embodiment, the dataset is constructed as follows: data on the retention of standardized solutes and organic micropollutants by nanofiltration membranes are obtained through literature review or experimental measurement; property characteristic data of organic micropollutants are obtained through public chemical databases; the above data are preprocessed to construct the dataset; the preprocessing includes deduplication, mean filling, or missing value processing.
[0057] In one embodiment, the input features of the machine learning model include: nanofiltration membrane retention data of standardized solutes; property characteristic data of organic micropollutants; and the output of the machine learning model is nanofiltration membrane retention data of organic micropollutants.
[0058] In one implementation, the machine learning model is an ensemble learning model based on decision trees, and the training process includes dataset partitioning, model training and performance evaluation, and parameter optimization to improve the model's prediction accuracy.
[0059] In one implementation, the retention data of the nanofiltration membrane to be tested for standardized solutes are tested; the retention data of the nanofiltration membrane to be tested for standardized solutes and the property characteristics data of the target pollutant are input into the mapping model; and the retention performance prediction results of the target organic micropollutant are output.
[0060] Example 1
[0061] A method for predicting the ability of nanofiltration membranes to remove organic micropollutants based on standardized solute retention performance, such as Figure 2 As shown, it includes the following steps:
[0062] The study collected data on the rejection rates of nanofiltration membranes for NaCl, Na2SO4, MgSO4, CaCl2, and MgCl2, as well as their rejection rates for 206 organic micropollutants, from literature. Molecular weight, charge state at pH 7, and log D value of these organic micropollutants were also collected from publicly available chemical information databases and prediction databases. These organic micropollutants encompassed categories such as perfluorinated and polyfluoroalkyl substances (PFASs), pharmaceutical and personal care products (PPCPs), disinfection byproducts (DBPs), endocrine disruptors (EDCs), and pesticides (PCs). For missing data points regarding the rejection rates of NaCl, Na2SO4, MgSO4, CaCl2, and MgCl2 from commercial nanofiltration membranes, the average value of the same commercial nanofiltration membrane was used for filling. A multi-random seeding strategy was employed to randomly divide the dataset into training and test sets, with a training-to-test set ratio of 8:2.
[0063] The Optuna framework was used to automatically optimize the key hyperparameters of the XGBoost model. The optimized parameters mainly included: maximum tree depth (max_depth), learning rate (learning_rate), number of trees (n_estimators), subsample ratio (subsample), feature sampling ratio (colsample_bytree), minimum leaf node weight (min_child_weight), and regularization parameters (reg_alpha, reg_lambda). During the optimization process, 40 hyperparameter search experiments were conducted, and 5-fold cross-validation was used to evaluate candidate parameter combinations in each experiment.
[0064] like Figure 3 As shown, most of the sample points in the training and test sets are distributed near the diagonal (the solid line at the diagonal represents the ideal prediction state, i.e., the predicted value equals the actual value), indicating that the model of this invention can reflect the correspondence between the single salt rejection performance and the rejection rate of organic micropollutants well. Especially in the medium-to-high rejection rate range, the consistency between the predicted and measured values is good, and the overall trend is relatively clear. From the performance indicators, the model's coefficient of determination R on the training set is [value missing]. 2 The R value was 0.93, and the RMSE was 6.6%; in the test set, the R value was... 2 The rejection ratio was 0.69, and the RMSE was 13.3%. These results indicate that the rejection rates of five standardized solutes (NaCl, Na₂SO₄, MgSO₄, CaCl₂, and MgCl₂) and the basic physicochemical properties of organic micropollutants can be used to effectively predict the retention behavior of nanofiltration membranes. While there are some differences between the training and test sets, the test set maintains a good correlation, indicating that the model constructed in this embodiment has a certain generalization ability.
[0065] Furthermore, to examine the applicability of the above model to different types of organic micropollutants, the prediction results were classified and statistically analyzed based on the molecular weight, charge state, and log D of the organic micropollutants in the test set. Specifically, organic micropollutants were divided into three ranges according to molecular weight: 0-200, 200-400, and greater than 400, to characterize organic micropollutants with different molecular size ranges; according to charge state, organic micropollutants were divided into three categories: negatively charged molecules (< -0.1), neutral molecules (-0.1-0.1), and positively charged molecules (> 0.1); and according to log D value, organic micropollutants were divided into three categories: hydrophilic molecules (< 0), molecules that are neither hydrophilic nor hydrophobic (0-3), and hydrophobic molecules (> 3).
[0066] Through the above classification and statistics, the predictive performance of the model of this invention in different types of organic micropollutant systems can be further compared, and the results are as follows: Figure 4As shown in the figure, the model prediction error was lowest (RMSE 10.46%) when the molecular weight of organic micropollutants was greater than 400 Da, indicating that the model's prediction ability for large-molecule organic micropollutants was more stable. This is because large-molecule organic micropollutants generally have a higher retention rate during nanofiltration, exhibiting a more pronounced size sieving effect, resulting in relatively concentrated variation patterns and easier model identification. In contrast, low-molecular-weight organic micropollutants are more likely to be affected by multiple factors during retention, leading to more dispersed prediction results. Based on the charge state grouping results, the prediction error for charged organic micropollutants was significantly lower than that for neutral pollutants. Among them, negatively charged organic micropollutants had the lowest RMSE at only 9.67%; positively charged organic micropollutants showed a slight increase in RMSE at 10.75%; while near-neutral organic micropollutants had the highest RMSE at 17.53%. This result indicates that the model is more sensitive to systems with strong charge effects. Since nanofiltration membranes are typically negatively charged, the retention of charged organic micropollutants is often subject to more pronounced electrostatic interactions, resulting in a clearer retention pattern. Conversely, neutral organic micropollutants lack a significant electrical response, and their retention is more easily affected by differences in molecular structure and interfacial interactions, making them relatively more difficult for the model to accurately characterize. In the classification results based on log D values, the model had the highest prediction error (RMSE of 15.47%) when log D > 3, exceeding the average level of the test set, indicating that the model's applicability to strongly hydrophobic pollutants is relatively weak. Unlike size sieving and electrostatic interactions, hydrophobic interactions involve more processes such as membrane surface adsorption and interfacial partitioning, information that is difficult to fully reflect through single-salt retention characteristics. Therefore, as the hydrophobicity of organic micropollutants increases, the model error also increases. Overall, the model performs well in pollutant systems dominated by size sieving and electrostatic interactions, but its predictive ability is limited in systems with more significant hydrophobic interactions. Therefore, it is more suitable for predicting nanofiltration retention dominated by size sieving and electrostatic interactions, while its prediction accuracy still has room for improvement for pollutants with stronger hydrophobicity and more complex mechanisms.
[0067] To further analyze the internal decision-making basis of the above model and quantitatively characterize the contribution of each input feature to the prediction results of organic micropollutant retention performance, the SHAP method is used for interpretation. Figure 5The SHAP importance distribution plot shows that the contributions of each input feature to the model output differ significantly. The molecular weight of organic micropollutants has the highest importance and the widest distribution range of SHAP values; higher molecular weights correspond to more positive SHAP values, while lower molecular weights are more prevalent in the negative region. This indicates that the larger the molecular weight of organic micropollutants, the more likely it is to improve the model's predicted rejection rate. In other words, size information remains the primary basis for the model's judgment of organic micropollutant rejection behavior, consistent with the understanding that nanofiltration is dominated by size sieving. Besides molecular weight, the rejection data of multiple single salts also rank highly in the importance ranking, indicating that the model does not rely solely on the properties of the organic micropollutants themselves for prediction, but simultaneously utilizes multiple single salt rejection information to characterize the separation characteristics of the nanofiltration membrane. Among these, the trend of NaCl rejection characteristics is relatively clear: low NaCl rejection rates correspond to more positive SHAP values, while high NaCl rejection rates show more negative contributions. In the NaCl system, ions have lower valence states and relatively weaker electrostatic interactions. Their retention behavior primarily reflects the membrane's basic sieving ability for small-sized solutes, and this, along with molecular weight, indicates that the size sieving effect dominates the model's decision-making process. Furthermore, the CaCl2 retention characteristic also exhibits some directionality: higher CaCl2 retention rates generally correspond to more positive SHAP values, while lower retention rates show more negative contributions, indicating that this characteristic can also provide information related to membrane separation capability. Although its distribution has some directionality, its overall importance ranking is relatively low. This may be because, on the one hand, the separation information reflected by CaCl2 overlaps to some extent with the characteristics of other single salts; on the other hand, it differs from NaCl in that it... + In comparison, Ca 2+ It has a larger hydration radius and a higher ionic valence state, and its effect may be influenced by multiple factors such as membrane sieving capacity and surface charge. However, for most samples, its influence is not as significant as that of Na. + Their contribution is more concentrated in a subset of samples.
[0068] In contrast, while other single-salt retention features are equally important, their directionality in the SHAP distribution is not as intuitive as that of NaCl and CaCl2, nor do they exhibit a consistent and clear monotonic trend. This indicates that these features do not correspond to a single mechanism in the model, but rather serve as supplementary information in the model's decision-making. Considering that the retention behavior of different salts is simultaneously influenced by multiple factors such as membrane pore size, charge characteristics, ion valence state, and mass transfer behavior, their physical meaning is inherently more complex, thus they do not exhibit a completely consistent regular distribution in the SHAP diagram. The above results demonstrate that the model of this invention does not simply equate single-salt retention features with a single mechanism variable, but rather comprehensively utilizes information from different salts to characterize the membrane's overall separation response to the ion system.
[0069] Figure 5 The results show that charge state is also of high importance, indicating that the electrical effect still plays a significant role in model decision-making. The SHAP distributions for different single salt characteristics are not entirely consistent; a notable difference is that the global importance of sulfate characteristics is generally higher than that of chloride characteristics, indicating that the model can distinguish the separation responses corresponding to different ion systems. Na₂SO₄ and MgSO₄ tend to correspond to positive SHAP contributions under higher eigenvalue conditions, suggesting that increased sulfate rejection is usually associated with higher predicted rejection rates. Considering that nanofiltration membranes are typically negatively charged under neutral conditions, divalent anion systems are more susceptible to the influence of membrane surface charge and Donnan repulsion. Therefore, sulfate rejection behavior often reflects membrane surface electrical characteristics more sensitively than chloride rejection. In other words, this difference is not merely a statistical difference in the salts themselves, but more likely a manifestation of membrane surface charge, ion valence state, and electrostatic repulsion behavior in the model. Therefore, the model's identification of electrical effects comes not only from the charge state of organic micropollutants themselves, but also from the membrane-ion interaction information carried by single salt rejection characteristics. In contrast, log D is relatively less important, with its SHAP values mostly concentrated near zero, showing only significant positive or negative contributions in a few samples. This indicates that hydrophobic interactions are not the dominant factor in the model of this invention, at least not as stably captured as size sieving effects and electrostatic interactions.
[0070] Overall, the conclusions of the SHAP analysis are relatively clear: the dominant factors identified by the model in this invention remain size sieving effect and electrostatic interaction, which is consistent with the general mechanism of nanofiltration separation of trace organic pollutants. More importantly, without explicitly introducing membrane physicochemical properties such as membrane pore size and zeta potential, the model in this invention can still reflect the key separation properties of the membrane using single-salt retention characteristics. This indicates that the performance-performance model constructed based on single-salt retention characteristics in this invention not only has good predictive ability but also possesses a certain mechanistic explanatory basis.
[0071] Furthermore, this invention also investigated the correlation between single salt rejection rate and the physicochemical properties of nanofiltration membranes. Considering the diverse sources of cross-document data, the potential deviation of variable distribution from normality, and the fact that variables may not necessarily satisfy a strict linear relationship, Spearman's rank correlation coefficient was used for analysis here. Figure 6The results show that the physical meanings corresponding to the retention characteristics of different single salts are not entirely the same. Overall, the chloride salt system exhibits high correlation, with CaCl2 and MgCl2 showing a correlation coefficient of 0.97, and NaCl with CaCl2 and MgCl2 reaching 0.83 and 0.80 respectively. This indicates that these salts largely share common information such as membrane structure compactness and surface electrical properties. Na2SO4 and MgSO4 also show a high correlation, with a correlation coefficient of 0.87, indicating strong consistency within the sulfate system. In contrast, the correlation between different salt systems containing the same cation is relatively limited. For example, the correlation coefficients between NaCl and Na2SO4, and between MgCl2 and MgSO4, are only 0.35 and 0.39 respectively. This suggests that even with the same cation, changes in the type and valence state of the anion can significantly affect the separation information reflected by the retention behavior. In other words, in comparing the characteristics of different single salts, the type of anion may be the more significant factor causing differences in characterization.
[0072] Analysis of the physicochemical properties of the membrane revealed differences in the responses of different single salts to key nanofiltration membrane attributes. Regarding structural parameters, the screening capacity (MWCO) of the nanofiltration membrane showed a significant negative correlation with the rejection rates of various single salts. For example, the correlation coefficients between MWCO and the rejection rates of NaCl, CaCl2, and MgCl2 were -0.68, -0.71, and -0.72, respectively. This indicates a good correlation between the size of the chloride salt system and the membrane. In terms of mass transfer characteristics, the pure water permeability (PWP) of the nanofiltration membrane also showed a negative correlation with the single salt rejection rate. For example, the correlation coefficients between PWP and the rejection rates of NaCl, CaCl2, and MgCl2 were -0.56, -0.64, and -0.65, respectively. Based on the general principles of nanofiltration membrane separation, this suggests that membranes with higher pure water permeability typically correspond to relatively lower salt rejection rates, reflecting a certain permeability-selectivity trade-off. Meanwhile, no significant and consistent correlation was observed between the water contact angle (WCA) and the rejection rates of individual salts, indicating that the ability of individual salt rejection characteristics to characterize the hydrophobicity of the membrane surface is relatively limited. In other words, individual salt rejection reflects more the membrane structure and interfacial electrical properties than the more complex hydrophobic interactions between the membrane and organic micropollutants.
[0073] In terms of electrical properties, the zeta potential of the membrane surface shows a more significant correlation with the rejection rates of some polyvalent salts. For example, the correlation coefficients between zeta potential and the rejection rates of Na₂SO₄ and MgSO₄ are 0.53 and 0.43, respectively, while its correlation with the chloride system is weaker, with an overall correlation coefficient close to 0, indicating no significant statistical association. This suggests that in the chloride system, size sieving and structural effects are more dominant than interfacial electrical properties. In contrast, sulfate rejection is not only constrained by the membrane structure but also more sensitive to the membrane surface charge, thus better reflecting the interfacial electrical characteristics of the membrane. For this reason, sulfate characteristics in the model are more likely to carry both nanofiltration membrane structural and electrical information simultaneously, rather than simply corresponding to a single mechanism variable. This also echoes the SHAP results mentioned earlier, that is, although sulfate characteristics are highly important, their contribution trend is not as strong as that of NaCl. In summary, although both chloride and sulfate can characterize membrane separation features, their focuses are not entirely consistent. Therefore, the significance of single-salt retention characteristics lies not in their correspondence to a single mechanism, but in their ability to characterize the comprehensive separation response of nanofiltration membranes to ion systems through the combination of information from different salts. There is some overlap in information among different single-salt characteristics, while also preserving complementarity at the mechanistic level. Meanwhile, correlation results also indicate that single-salt retention characteristics primarily reflect the membrane's structural features and interfacial electrical properties, while their coverage of membrane surface hydrophobicity information is relatively limited.
[0074] In summary, a relatively stable statistical correlation exists between the single-salt rejection rate and membrane pore structure, mass transfer characteristics, and interfacial electrical properties. This indicates that although single-salt rejection performance is not a physicochemical property parameter of the membrane itself, it can reflect the key separation attributes of the membrane to a certain extent and can serve as an indirect characterization parameter for membrane separation characteristics in performance-performance models. However, its characterization of hydrophobic interactions is still insufficient, therefore it needs to be supplemented by interfacial parameters such as WCA. Overall, using single-salt rejection characteristics as input variables to construct performance-performance prediction models has a clear statistical basis and relatively reasonable mechanistic support.
[0075] Based on this, the present invention compares the following four types of feature combination models: (1) using only NaCl as a standardized solute to characterize the membrane structure and mass transfer related information reflected by the chloride salt system; (2) further introducing Na2SO4 on the basis of NaCl to supplement the membrane surface electrical related information carried by the sulfate system; (3) using NaCl, Na2SO4, MgSO4, CaCl2 and MgCl2 as standardized solutes to examine the effect of multi-salt combination on the model prediction performance; (4) adding WCA on the basis of five standardized solutes (NaCl, Na2SO4, MgSO4, CaCl2, MgCl2) as a supplementary variable for the hydrophilic and hydrophobic characteristics of the membrane surface for comparative analysis. Figure 7As shown, when only NaCl is used as the standardized solute, the model prediction error is relatively high, with an RMSE of approximately 14.4%. Introducing Na₂SO₄ further reduces the RMSE to approximately 13.6%, indicating that adding sulfate-retention characteristics can provide membrane-related separation information beyond NaCl. Using five typical single salts as standardized solutes, the RMSE further decreases to 13.3%, R... 2 The RMS value was 0.69, but compared to the previous stage, the decrease in model error was significantly reduced. This result indicates that after the combined introduction of NaCl and Na2SO4, the model has captured the main information of the film structure and interfacial electrical properties. While adding other single salt features can still bring some gains, their supplementary effect is relatively limited, showing a trend of gradually weakening marginal contributions. Furthermore, after introducing WCA into the five single salts, the RMSE further decreased from approximately 13.3% to 11.7%, R... 2 Increased from 0.69 to 0.76 (see) Figure 8 This improvement is significantly higher than the marginal improvement brought about by further adding single salt features, indicating that the addition of information dimensions can effectively improve model performance. The introduction of WCA supplements the model with hydrophilic and hydrophobic information of the membrane surface, and can reflect the hydrophilic and hydrophobic matching relationship between the membrane and pollutants together with the log D feature of organic micropollutants, thereby improving the model's ability to predict the retention behavior of complex organic pollutants.
[0076] Figure 8 The results show that after adding WCA, the scattered points in both the training and test sets are still generally distributed along the diagonal, indicating that the model can reflect the correspondence between experimental and predicted values well. Compared with the model using only five single-salt cutoff features, the dispersion of the scattered points is significantly reduced after adding WCA, especially the test set scattered points, which are more concentrated near the diagonal. In the medium-to-high cutoff range, the deviation of sample points from the diagonal is reduced, and some samples with large deviations are also improved. At the same time, the clustering trend of samples in the high cutoff region along the diagonal is more obvious, indicating that the model's predictive stability for samples in this range is improved. This verifies the supplementary role of WCA in the performance-performance framework constructed by the single-salt cutoff feature.
[0077] In summary, compared with the property-to-performance models commonly used in existing technologies, the RMSE of the model constructed based on performance-to-performance in this invention is in a better range and close to the good prediction level reported in the literature (see Table 1). This indicates that the model can still achieve good prediction results by relying on only a few relatively easy-to-obtain performance features and adding the WCA interface parameter. It is particularly noteworthy that the model of this invention does not rely on explicit characterization parameters such as membrane pore size and zeta potential, which are difficult to obtain uniformly. It achieves a good balance between operability and predictive ability, making it more conducive to rapid prediction and method promotion for practical separation tasks.
[0078] Table 1: Comparison of the performance of machine learning models in predicting the retention of trace organic pollutants by nanofiltration membranes in recent years.
[0079]
[0080] The models in Table 1, namely DD model with DLM, DD model without DLM, DKD model with DLM, and DKD model without DLM, are derived from the literature: Wang H. et al., Understanding Rejection Mechanisms of Trace Organic Contaminants by Polyamide Membranes via Data-Knowledge Codriven Machine Learning, *Environmental Science & Technology*, 2024, Vol. 58, No. 13.
[0081] The models XGBoost-CAMES and XGBoost-TPE in Table 1 are from the literature: Hu A. et al., A machine learning based framework to tailor properties of nanofiltration and reverseosmosis membranes for targeted removal of organic micropollutants, Water Research, 2025, Vol. 268, Part A;
[0082] The XGBoost model in Table 1 is from the literature: Jeong N. et al., Elucidating governing factors of PFAS removal by polyamide membranes using machine learning and molecular simulations, Nature Communications, 2024, Vol. 15;
[0083] The models in Table 1, namely Single MF-model, D-model, and DMF-MRL model, are from the literature: Lu D. et al., Asmart framework to design membranes for organic micropollutants removal, Nature Sustainability, 2025, Vol. 8, No. 10, pp. 1177-1189.
[0084] The XGBoost model in Table 1 is derived from the literature: Wang H. et al., Customizing interpretable machine learning strategy to unravel causal mechanisms in organic micropollutant rejection by polyamide membranes, Journal of Membrane Science, 2026, Vol. 738, Part A.
[0085] The implementation of the above embodiments of the present invention is based on programmed processing by a device with processor functionality. Therefore, in engineering practice, the technical solutions and functions of the various embodiments of the present invention are encapsulated into various modules. Based on this reality, and on the basis of the above embodiments, the embodiments of the present invention provide a prediction system for the retention performance of nanofiltration membranes for trace organic matter, used to implement the above-mentioned method for predicting the retention performance of nanofiltration membranes for trace organic matter. The system includes: a data acquisition module: used to acquire retention data of standardized solutes and organic micropollutants by the nanofiltration membrane, and to acquire property characteristic data of organic micropollutants, and to construct a dataset; a model building module: used to train a machine learning model using the dataset, and to establish a mapping model for predicting the retention performance of organic micropollutants by the standardized solute retention performance; and a retention performance prediction module: used to predict the retention performance of the nanofiltration membrane under test for target organic micropollutants using the mapping model.
[0086] The method described in this invention relies on electronic devices; therefore, it is necessary to introduce the relevant electronic devices. To this end, this invention provides a device for predicting the retention performance of nanofiltration membranes for trace organic matter. This device includes: one or more processors; and a storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method for predicting the retention performance of nanofiltration membranes for trace organic matter. The storage device can be any non-volatile storage device such as a hard disk, solid-state drive, flash drive, or optical disk, used to store computer programs and necessary data files. The stored computer programs include: a data acquisition module, a model building module, and a retention performance prediction module.
[0087] The above detailed embodiments describe the implementation of the present invention; however, the present invention is not limited to the specific details described in the above embodiments. Within the scope of the claims and technical concept of the present invention, various simple modifications and changes can be made to the technical solution of the present invention, and these simple modifications all fall within the protection scope of the present invention.
Claims
1. A method for predicting the retention performance of nanofiltration membranes for trace organic matter, characterized in that, Includes the following steps: Data on the retention of standardized solutes by nanofiltration membranes, data on the retention of organic micropollutants, and the property characteristics of the organic micropollutants are collected, and a dataset is constructed. The aforementioned dataset was used to train the machine learning model, and a mapping model was established to predict the retention performance of organic micropollutants based on the standardized solute retention performance. The retention data of the nanofiltration membrane under test for the standardized solute and the property characteristics data of the target organic micropollutant are input into the mapping model, thereby realizing the prediction of the retention performance of the nanofiltration membrane under test for the target organic micropollutant.
2. The method for predicting the retention performance of nanofiltration membranes for trace organic matter according to claim 1, characterized in that, The standardized solute includes at least one monovalent salt and at least one polyvalent salt.
3. The method for predicting the retention performance of nanofiltration membranes for trace organic matter according to claim 2, characterized in that, The monovalent salts include: NaCl, KCl; and / or, the polyvalent salts include: Na2SO4, MgSO4, CaCl2, MgCl2.
4. The method for predicting the retention performance of nanofiltration membranes for trace organic matter according to claim 2, characterized in that, The standardized solute also includes: at least one small organic molecule with a certain degree of hydrophobicity and / or at least one neutral small molecule.
5. The method for predicting the retention performance of nanofiltration membranes for trace organic matter according to claim 4, characterized in that, The organic small molecules with a certain degree of hydrophobicity refer to organic small molecules with log D greater than 3, including: carbamazepine, bisphenol A, ibuprofen, diclofenac, atrazine; and / or, the neutral small molecules include: glucose, polyethylene glycol, sucrose.
6. The method for predicting the retention performance of nanofiltration membranes for trace organic matter according to claim 4, characterized in that, For a positively charged nanofiltration membrane to be tested, the standardized solute comprises: at least one monovalent salt and at least one polyvalent salt containing a divalent cation; and / or, for a negatively charged nanofiltration membrane to be tested, the standardized solute comprises: at least one monovalent salt and at least one polyvalent salt containing a divalent anion; and / or, for a charged target organic micropollutant, the standardized solute comprises: one monovalent salt and one polyvalent salt; for a hydrophobic target organic micropollutant, the standardized solute comprises: at least one monovalent salt, at least one polyvalent salt, and at least one small organic molecule with a certain degree of hydrophobicity.
7. The method for predicting the retention performance of nanofiltration membranes for trace organic matter according to claim 1, characterized in that, The property characteristics data of the organic micropollutants include: molecular weight, and the logarithmic values of charge state and partition coefficient under pH = 7 conditions.
8. The method for predicting the retention performance of nanofiltration membranes for trace organic matter according to claim 1, characterized in that, The dataset is constructed as follows: data on the retention of standardized solutes and organic micropollutants by nanofiltration membranes are obtained through literature review or experimental measurement; property characteristic data of organic micropollutants are obtained through public chemical databases; the above data are preprocessed to construct the dataset; the preprocessing includes deduplication, mean filling or missing value processing.
9. A system for predicting the retention performance of nanofiltration membranes for trace organic matter, used to implement the method for predicting the retention performance of nanofiltration membranes for trace organic matter as described in any one of claims 1-8, characterized in that, include: Data acquisition module: used to acquire data on the retention of standardized solutes and organic micropollutants by nanofiltration membranes, as well as the property characteristics data of organic micropollutants, and to construct a dataset; Model building module: used to train the machine learning model using the dataset and establish a mapping model that predicts the retention performance of organic micropollutants based on the standardized solute retention performance; Retention performance prediction module: Used to predict the retention performance of the nanofiltration membrane under test for target organic micropollutants using a mapping model.
10. A device for predicting the retention performance of nanofiltration membranes for trace organic matter, characterized in that, include: One or more processors; And a storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the method for predicting the retention performance of a nanofiltration membrane for trace organic matter as described in any one of claims 1-8.