Method for predicting separation performance of mixed matrix membrane through machine learning

The gas separation performance prediction model of hybrid matrix membrane is constructed through machine learning methods, which solves the problem of time-consuming and costly design and manufacturing of MMMs in the prior art, and achieves efficient and accurate MMMs performance prediction and design.

CN119993333APending Publication Date: 2025-05-13NANJING TECH UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411824521.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-11
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art faces the problem of the "trial-error" phase when designing and manufacturing hybrid matrix membranes (MMMs), resulting in time-consuming, labor-intensive, cost-effective, and lack of effective machine learning methods to predict the gas separation performance of MMMs.

Method used

The gas separation performance prediction model of mixed matrix membranes is constructed through machine learning methods, and the machine learning model is trained to predict the gas separation performance of MMMs using characteristic variables such as the physical properties parameters of polymers and the structural characteristics of fillers.

Benefits of technology

The gas separation performance of the mixed matrix membrane was successfully predicted, which improved the efficiency and accuracy of MMMs design, reduced development costs, and provided an explanation of the structure-performance relationship of MMMs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993333A_ABST
    Figure CN119993333A_ABST
Patent Text Reader

Abstract

The invention relates to a method for predicting the separation performance of a mixed matrix membrane through machine learning, and belongs to the technical field of computational material chemistry. Due to the reasons of various filler types, complex membrane structures, indefinite mass transfer mechanisms and the like, the design and preparation of mixed matrix membranes (MMMs) are still in a trial and error stage. The invention discloses a high-throughput design and screening mixed matrix membrane based on machine learning acceleration. Through high-throughput giant regular Monte Carlo calculation and molecular dynamics simulation, the reliable structure and performance characteristics of the MOFs filler are calculated. The high-throughput calculation data and the experimental data of the polymer membrane are combined to construct a large data set. A machine learning regression model is trained to predict penetration and separation performance, and a potential structure-performance relationship of MMMs is disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a method for predicting separation performance of a mixed matrix membrane through machine learning, and belongs to the technical field of computational material chemistry. Background Art

[0002] Mixed matrix membranes (MMMs) exhibit excellent separation performance by dispersing inorganic fillers with specific selectivity in a polymer matrix, and have great potential to break through the Robeson upper limit. Metal-organic frameworks (MOFs), as a class of materials with highly ordered pore structures and tunable chemical functionality, are considered to be ideal filler candidates for MMMs. In recent years, exciting progress has been made in the research on MOF mixed matrix membranes. Despite the great progress and breakthroughs in separation performance, the design and fabrication of MMMs are still in the "trial and error" stage due to the theoretically unlimited types of fillers, complex membrane structures, and unclear mass transfer mechanisms, which is not only time-consuming, labor-intensive, but also costly.

[0003] Currently, most studies involving machine learning in the separation field focus on the screening of MOF membranes and polymer membranes. In order to expand the application of ML in the design and screening of MMMs, some major challenges still need to be overcome: (1) The experimental separation data of many gas pairs in MMMs are insufficient to train ML models; therefore, there is an urgent need to develop methods to construct MMMs datasets. (2) Most ML studies initially predict the properties of MOFs and use them as a basis to calculate the properties of MMMs. However, direct ML studies on MMMs datasets to decipher the structure-property relationship of MMMs are still quite scarce.

[0004] Helium is a non-renewable resource and is relatively scarce in nature. Currently, extracting helium from natural gas is the only feasible way to utilize helium, and this process consumes a lot of energy. Among the many gas separation technologies, membrane separation technology has attracted widespread attention for its advantages such as low energy consumption, simple operation, and easy scalability. Summary of the invention

[0005] The present invention provides a method for constructing the gas separation performance of a mixed matrix membrane by a machine learning method, and the method successfully obtains the prediction of the gas separation performance of the mixed matrix membrane by characteristic variables thereof.

[0006] The technical solution is:

[0007] A method for predicting the gas separation performance of a mixed matrix membrane by machine learning comprises the following steps:

[0008] Step 1, setting the physical property parameters of the mixed matrix membrane as input variables and setting the gas separation performance of the mixed matrix membrane as target variables;

[0009] Step 2, using the data of the training samples to train the machine learning model to obtain a prediction model;

[0010] Step 3: Use the prediction model to predict the gas separation performance of the mixed matrix membrane.

[0011] The mixed matrix membrane refers to a membrane obtained by mixing a polymer and a filler and capable of separating gases.

[0012] The filler is one of MOF material, COF material, two-dimensional sheet material, molecular sieve, carbon nanotube, graphene, metal oxide, transition metal sulfide and the like.

[0013] The polymer is selected from polysulfone (PSU), polyimide (PI), polyethersulfone (PES), polyetherimide (PEI), polycarbonate (PC), polytetrafluoroethylene (PTFE), polyvinylidenefluoride (PVDF), polyvinyl alcohol (PVA), polyacrylonitrile (PAN), polylactic acid (PLA), polystyrene (PS), polyvinyl chloride (PVC), polypropylene (PP), polyethylene terephthalate (PET), polyurethane (PU), etc.

[0014] The mass percentage of the filler in the polymer is in the range of 0.01-20%.

[0015] The gas separation refers to CO2, CH4, N2, O2, H2, He, CO, H2O, NH3, H2S, SO2, NO x , CH3OH, C3H8, C2H6, CO or a mixture of two or more gases can be separated.

[0016] The machine learning model is selected from one or a combination of decision tree (DT), random forest (RF), lightweight gradient boosting machine (LightGBM) and extreme gradient boosting (XGBoost), support vector machine (SVM), logistic regression (LogisticRegression), naive Bayes (NaiveeBayes), K-Nearest Neighbors (KNN), linear regression (LinearRegression), neural networks (NeuralNetworks), deep learning (DeepLearning), convolutional neural networks (ConvolutionalNeuralNetworks, CNN), recurrent neural networks (RecurrentNeuralNetworks, RNN), long short-term memory network (LongShort-TermMemory, LSTM), clustering algorithm (such as K-Means), principal component analysis (PCA), natural language processing model (such as BERT), generative adversarial networks (GenerativeAdversarialNetworks, GANs); decision tree (DT), random forest (RF), lightweight gradient boosting machine (LightGBM) and extreme gradient boosting (XGBoost) are preferred; extreme gradient boosting (XGBoost) is further preferred.

[0017] The physical property parameters include: polymer permeability to one component, polymer permeability to another component, polymer selectivity to two components, temperature, filler volume fraction in the mixed matrix membrane, filler global maximum pore diameter Maximum cavity diameter of packing Packing pore size limit diameter Filler density ρ (g / cm3), filler accessible specific surface area VSA (m2 / cm3), filler accessible pore volume V p (cm3 / g) and one or more combinations of filler accessible porosity Φ, filler type, and polymer type; preferably, a combination of the permeability of the polymer to one component, the permeability of the polymer to another component, the selectivity of the polymer to the two components, the volume fraction of the filler in the mixed matrix membrane, the filler pore limiting diameter PLD, the filler accessible porosity Φ, the filler type, and the polymer type.

[0018] The gas separation performance refers to one or a combination of the permeability of the mixed matrix membrane to one gas, the separation selectivity of the mixed matrix membrane to two gas components.

[0019] The permeability of the mixed matrix membrane to one gas is calculated by the following formula:

[0020]

[0021] P i 0 is the permeability of the filler to gas, is the volume fraction of filler and interfacial voids, represents the permeability of the polymer, P i represents the permeability of MMM to gas;

[0022] And, P i 0 It is calculated by the following formula:

[0023] P i 0 =K i 0 ×D i 0

[0024] P i 0 is the permeability coefficient of gas i, K i 0 is the Henry constant for gas i, D i 0 is the gas i diffusion coefficient.

[0025] Henry's constant is calculated by the Widom particle insertion method under infinite dilution conditions of gas, and the diffusion coefficient is calculated from the slope of the linear region of the mean square displacement of gas molecules versus time.

[0026] The separation selectivity of the mixed matrix membrane for two gas components is calculated by the following formula:

[0027]

[0028] P represents the permeability coefficient of the gas in the filler, where i and j represent different gas molecules. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 : Flowchart of strategies for building a large dataset of MMMs and conducting machine learning research on MMMs datasets.

[0030] Figure 2 : Preliminary analysis of the MMMs dataset. (a) Diversity of input features and (b, c, d) scatter plots of membrane performance versus input features.

[0031] Figure 3 : Pearson correlation matrix analysis of input and target features for (a) the initial dataset and (b) the reduced dataset.

[0032] Figure 4 : Comparison of loss function results of four models.

[0033] Figure 5 :(a) Helium permeability of MMMs in training and test sets (P He ) and (b) helium / methane selectivity (S He / CH4 ) and the scatter plot of the predicted values ​​and the high-throughput calculation simulation values.

[0034] Figure 6 :XGBoost model prediction value and (a)P He and (b) S He / CH4 Comparison of calculated values. DETAILED DESCRIPTION

[0035] The method of the present invention is used to predict, explain and screen high performance mixed matrix membranes for helium separation (see Figure 1 ). The structural and performance characteristics of 10,143 experimentally synthesized MOFs from the CoRE MOF 2019 database were calculated by Grand Canonical Monte Carlo (GCMC) and Molecular Dynamics (MD) simulations. These high-throughput computational data were combined with experimental data of 26 polymers collected from the literature, and a dataset of 456,872 MMMs samples was constructed through the Maxwell model to develop a machine learning model. After a comprehensive evaluation, the extreme gradient boosting (XGBoost) model was shown to perform the best among the four trained ML models, and the model interpretation analysis of the XGBoost model revealed the structure-performance relationship of the MMMs. Subsequently, the XGBoost model was applied to new MMMs composed of computationally constructed hypothetical MOFs (hMOFs), and our ML model was able to accurately predict the helium separation performance of hMOFs / polymer MMMs.

[0036] We developed a high-throughput computer simulation (HTCS) method to calculate the gas permeability and selectivity of helium and methane in MMMs. All structural properties of MOFs involved, including the global maximum pore diameter (GCD, ), maximum cavity diameter (LCD, ), aperture limiting diameter (PLD, ), density (ρ, g / cm 3 ), accessible specific surface area (VSA, m 2 / m 3 ), accessible pore volume (V p ,cm 3 / g) and accessible porosity (Φ) were calculated using Zeo++ software.

[0037] use The probe radius (representing the radius of the nitrogen molecule) is used to predict the accessible specific surface area of ​​the studied material.

[0038] The datasets in this patent are as follows:

[0039] Initially, the structures of 10,143 ordered MOFs from the 2019 CoRE (computation-ready, experimentally based) MOF database were calculated. In this database, the structural parameters of MOFs are derived from experimental data, and the solvent and ligand molecules in the framework have been removed. [1] The crystal structures of the CoRE MOF database can be directly used for molecular simulations without further modification, which is a good template for exploring ML strategies to design MOFs and MOF-based materials. Then, 9,008 MOFs with PLDs between 2.6 and 10 Å were selected. Finally, the atomic charges of these MOFs were calculated using the PACMOF (Partial Atomic Charges in Metal-Organic Frameworks) code. [5] PACMOF charges have been successfully used to describe Coulomb interactions in previous studies. [6-8] In addition, 221 MOFs had structural disorder or errors at their metal sites and their atomic charges could not be calculated, so 8,787 MOFs were left for further study.

[0040] Non-patent literature 1-8:

[0041] YGChung,E.Haldoupis,BJBucior,M.Haranczyk,S.Lee,H.Zhang,KDVogiatzis,M.Milisavljevic,S.Ling,JSCamp,B.Slater,JISiepmann,DSSholl,RQSnurr,Advances,Updates,and Analytics for the Computation-Ready,Experimental Metal–Organic Framework Database:CoRE MOF 2019,Journal ofChemical&Engineering Data 64(12)(2019)5985-5998.https: / / doi.org / 10.1021 / acs.jced.9b00835.

[0042] R.Anderson,A.Biong,D.A.Gómez-Gualdrón,Adsorption Isotherm Predictionsfor Multiple Molecules in MOFs Using the Same Deep Learning Model,Journal ofChemical Theory and Computation 16(2)(2020)1271-1283.https: / / doi.org / 10.1021 / acs.jctc.9b00940.

[0043] H.Daglar,H.C.Gulbalkan,G.Avci,G.O.Aksu,O.F.Altundal,C.Altintas,I.Erucar,S.Keskin,Effect of Metal-Organic Framework(MOF)Database Selection onthe Assessment of Gas Storage and Separation Potentials of MOFs,AngewandteChemie-International Edition 60(14)(2021)7828-7837.https: / / doi.org / 10.1002 / anie.202015250.

[0044] A.Nandy,C.R.Duan,H.J.Kulik,Using Machine Learning and Data Mining toLeverage Community Knowledge for the Engineering of Stable Metal-OrganicFrameworks,Journal of the American Chemical Society 143(42)(2021)17535-17547.https: / / doi.org / 10.1021 / jacs.1c07217.

[0045] S.Kancharlapalli,A.Gopalan,M.Haranczyk,R.Q.Snurr,Fast and AccurateMachine Learning Strategy for Calculating Partial Atomic Charges in Metal-Organic Frameworks,Journal of Chemical Theory and Computation 17(5)(2021)3052-3064.https: / / doi.org / 10.1021 / acs.jctc.0c01229.

[0046] J.Burner,J.Luo,A.White,A.Mirmiran,O.Kwon,P.G.Boyd,S.Maley,M.Gibaldi,S.Simrod,V.Ogden,T.K.Woo,ARC-MOF:A Diverse Database of Metal-OrganicFrameworks with DFT-Derived Partial Atomic Charges and Descriptors forMachine Learning,Chemistry of Materials(2023).https: / / doi.org / 10.1021 / acs.chemmater.2c02485.

[0047] R.Wang,B.C.Bukowski,J.X.Duan,J.Y.Sui,R.Q.Snurr,J.T.Hupp,Art ofArchitecture:Efficient Transport through Solvent-Filled Metal-OrganicFrameworks Regulated by Topology,Chemistry of Materials 33(17)(2021)6832-6840.https: / / doi.org / 10.1021 / acs.chemmater.1c01536.

[0048] H. Demir, S. Keskin, Zr-MOFs for CF4 / CH4, CH4 / H2, and CH4 / N2 separation: toward the goal of discovering stable and effective adsorbents, MolecularSystems Design&Engineering 6(8)(2021)627-642. https: / / doi.org / 10.1039 / d1me00060h.

[0049] Grand Canonical Monte Carlo (GCMC) simulations were performed at 308 K using the Widom particle insertion method to calculate the Henry constants (K) for helium and methane at infinite dilution in MOFs. 0 )[B. Widom, Some Topics in the Theory of Fluids, The Journal of Chemical Physics 39(11)(1963)2808-2812. https: / / doi.org / 10.1063 / 1.1734110.]. For each MOF, Monte Carlo (MC) simulations were run for 10 5 cycles, of which the first 50,000 cycles were used for equilibrium and the next 50,000 cycles were used for ensemble averaging. In the molecular dynamics (MD) simulations, we inserted 30 gas molecules in each MOF to simulate the diffusion behavior of He and CH4 gas molecules in MOF under ideal conditions. The mean square displacement (MSD) of the gas molecules was obtained at the end of the MD simulation. The diffusion coefficient is related to the MSD (t) of the labeled particles through the Einstein relation, using the dimensionality of the diffusion (d) at time t as follows:

[0050]

[0051] The mean square displacement (MSD) is proportional to time, and when the slope of this linear region is divided by 6 (for a three-dimensional system), the diffusion coefficient (D 0)[H.Daglar, S.Keskin, Recent advances, opportunities, and challenges in high-throughput computational screening of MOFs for gas separations, Coordination Chemistry Reviews 422(2020). https: / / doi.org / 10.1016 / j.ccr.2020.213470.]. Molecular dynamics (MD) simulations were performed in the NVT ensemble using Nose-Hoover with a simulation time step of 1 femtosecond and a total simulation time of 2 nanoseconds, of which the last 1 nanosecond was the production simulation. Considering D 0 He ≤10 -12 m 2 / s MOFs have very low separation efficiency, only D 0 He >10 -12 m 2 The MOFs with K / s were retained for further study, resulting in 8786 MOFs. 0 and D 0 The helium permeability (P) of each MOF was calculated according to equation (3). 0 He ) and methane permeability (P 0 CH4 ):

[0052] P 0 =K 0 ×D 0 (3)

[0053] Data on polymer membranes were collected from the literature in the Web of Science Core Collection database (Table 1) to construct a structure-property dataset of MOF / polymer MMMs.

[0054] Table 1 Based on the literature, 26 different polymer membranes in He / CH 4 Separation performance data

[0055]

[0056] The Maxwell model is the most widely used permeation model for obtaining the gas separation performance of MMMs, especially when the filler loading is low (volume fraction ≤ 20%). The permeability of the mixed matrix membrane for the gas separation process is calculated by the following formula.

[0057]

[0058] Where P i 0 is the permeability of the MOF calculated by molecular simulation, is the volume fraction of the dispersed phase (filler MOF) and the interfacial voids, Pi1 represents the permeability of the pure polymer, and P i represents the permeability of MMM to gas. An additional assumption is Roughly corresponds to the volume of the dispersed phase, ie the volume fraction of the MOF used as filler in the polymer matrix.

[0059] Filler loadings of 5 vol% and 20 vol% were selected To ensure that it is within the applicable range of the Maxwell model. The gas selectivity (S) of MMMs is calculated by equation (4): i / j , where i and j represent different gas molecules), a total of 456,872 MMMs samples were generated.

[0060]

[0061] Taking into account the complexity of the data applied to He separation membranes, we selected four typical machine learning models for prediction and tuning, namely decision time (DT), random forest (RF), lightweight gradient boosting machine (LightGBM) and extreme gradient boosting (XGBoost). Finally, we evaluated and selected the best model for application.

[0062] Ensemble learning methods such as decision trees, random forests, lightweight gradient boosting machines (LightGBM), and extreme gradient boosting (XGBoost) all rely on a series of parameters to optimize model performance. The model is built by calling the corresponding python toolkit. Specifically, the max_depth parameter of the decision tree defaults to None, allowing the tree depth to grow indefinitely until the leaf nodes are pure; min_samples_split and min_samples_leaf default to 2 and 1, respectively, controlling the number of samples for node splits and leaf nodes; max_features defaults to None, meaning that all features are considered when splitting. The n_estimators of the random forest defaults to 100, which enhances the stability of the model by building multiple decision trees; bootstrap defaults to True, allowing each tree to use bootstrap sampling. The num_leaves and max_depth of LightGBM default to 31 and -1, respectively, to adjust the complexity of the tree; learning_rate and n_estimators default to 0.1 and 100, controlling the learning step size and the number of trees. XGBoost's n_estimators and learning_rate are the same as LightGBM, with default values ​​of 100 and 0.1; max_depth defaults to 6, limiting the tree depth to prevent overfitting; colsample_bytree and min_child_weight default to 1 and 1 respectively, the former controls the feature sampling ratio, and the latter controls the minimum weight and split node. These parameters together determine the learning and generalization capabilities of the model.

[0063] In order to optimize and validate the model robustly, we used the random search cross-validation method to optimize the hyperparameters of the four regression models. This function randomly selects parameter combinations for model training and uses cross-validation (cv=10) to evaluate model performance: the training set is randomly divided into 10 different subsets, each subset is called a fold, and then the decision tree model is trained and evaluated 10 times - each time 1 fold is selected for evaluation and the other 9 folds are used for training. In machine learning, 80% of the dataset is randomly divided into training sets, and the remaining 20% ​​constitutes the test dataset. Model performance is evaluated by the loss functions RMSE and MAE (mean absolute error). In addition, the coefficient of determination, i.e. R2, is used to evaluate the goodness of fit of the model.

[0064] Validation of calculation method

[0065] The helium permeabilities of CuBTC, IRMOF-3, ZIF-68, ZIF-8, MMOF, and Cu2(bza)4(pyz) were calculated under the same experimental conditions as reported in the literature. The calculated and experimental results are in good agreement, as reflected by the high R2 value of 0.82. The good agreement between the experimental reports and our calculated results validates the rationality of the current force field and atomic charges. The gas permeabilities of MOF fillers obtained from high-throughput computational simulations are combined with those of polymer matrices collected from the literature using the Maxwell model to obtain the gas permeabilities of MMMs. To verify the accuracy of the calculation method for calculating the gas separation performance of MOF / polymer MMMs, we calculated the gas permeabilities of several MMMs under the same conditions reported in the literature. These MMMs consist of 5 different MOFs (including CuBTC, MIL-53, MOF-5, UiO-66, ZIF-8) and 12 different polymers (including Azide-PMP, Matrimid5218, 6FDA-DAM, PVC-g-POEM, XLPEGDA, 6FDA-durene, MH-1657, 6FDA-DAM, PIM-1, Pebax-2533, Ultem 1000). The gas permeability obtained by the calculation method is not generally overestimated compared with the experimental data.

[0066] Evaluation of Machine Learning

[0067] In the initial dataset, there are 14 variables, including 12 input variables and 2 target variables. The 10 numerical input variables are as follows: 1 He (Helium permeability of polymers, Barrer), P 1 CH4 (methane permeability of polymers, Barrer), S 1 He / CH4 (helium / methane selectivity of the polymer), T (temperature, °C), loading (volume fraction of MOFs in MMMs), ρ(g / cm3), VSA(m2 / cm3), Vp(cm3 / g), and Φ. The remaining two input variables are the types of MOFs and polymers, which are categorical variables and are digitized using One-Hot Encoding. The two target variables are P He (Helium permeability of MMM, Barrer) and S He / CH4 (Helium / methane selectivity of MMM). There are no missing data in the dataset.

[0068] Figure 2A preliminary analysis of the numerical input descriptors is shown. The helium and methane permeabilities and helium / methane selectivities for the 26 polymers range from 0 to 2300 barrers, 0 to 300 barrers, and 0 to 1000, respectively. The PLD for MOFs varies from 2.6 to 10 angstroms and Φ ranges from 0.2 to 0.9, indicating the diversity and complexity of the dataset. It is difficult to draw any conclusions from the permeability-selectivity scatter plots and given input variables such as MOF pore size, MOF porosity, and polymer properties due to the influence of many other parameters.

[0069] Feature selection was performed using Pearson correlation analysis. Figure 3 We can see that there are strong correlations between some of the input features, which means that the current dataset is redundant. Therefore, some of these features are removed to simplify the dataset. The new dataset includes 6 input features: 1 He , P 1 CH4 , S 1 He / CH4 、loading、 and Φ, which are weakly correlated with each other but strongly correlated with the target feature. Removing redundant features can greatly reduce the complexity of the ML model, save a lot of computing resources, and enhance the interpretability of the model; in addition, two more feature variables, the type of MOF material and the type of polymer, were added, with a total of 8 feature values ​​as input variables. After feature selection, the dataset was normalized and then used to develop a P-based predictor for MMMs. He and S He / CH4 ML models. By using the random search cross-validation method, we found the best parameter combination in a large parameter space to improve the prediction performance of the model. As the number of samples increased, the RMSE of all four ML models on the training and validation datasets gradually converged, with small differences, meaning that all models avoided overfitting.

[0070] Figure 4 Shows the P on the test dataset He and S He / CH4 The R2, RMSE and MAE results show that all four models have good He and S He / CH4 All showed high goodness of fit (>0.95). He and S He / CH4 The XGBoost model has the lowest RMSE value. He and S He / CH4 The MAE value of the XGBoost model is also competitive among the four ML models. Therefore, the XGBoost model is selected for further research.He and S He / CH4 The scatter plots of HTCS values ​​in the training and test datasets are shown in Figure 5 In (a) and (b) of the training data set, P He and S He / CH4 The R2 values ​​in the test dataset are 0.9938 and 0.9916, respectively, and 0.9929 and 0.9866, respectively, which are quite outstanding results given the diversity of the dataset and the inherent complexity of the machine learning algorithm. RMSE and MAE are related to the range of the target feature, and according to their mathematical expressions, RMSE is more sensitive than MAE. In general, the larger the range of the target feature, the higher the RMSE and MAE values. Considering P He (0 to 4000 Barrer) and S He / CH4 Achieving RMSE and MAE values ​​within 100 in both training and testing datasets indicates that the model has strong generalization ability, and the close similarity of RMSE and MAE values ​​between the testing and training datasets further indicates that the model is not overfitting.

[0071] An important advantage of machine learning models is that these models can be transferred to new, unexplored material spaces, making accurate predictions for previously unseen materials. We studied the transferability of the resulting XGBoost models on new MMMs constructed from different MOFs from the experimental CoRE MOF dataset. The MOFs in the new MMMs are computationally constructed MOFs from the hMOF database [CE Wilmer, M. Leaf, CY Lee, OK Farha, BG Hauser, J T Hupp, RQ Snurr, Large-scale screening of hypothetical metal-organic frameworks, Nature Chemistry 4(2) (2012) 83-89. https: / / doi.org / 10.1038 / nchem.1192.]. They were combined with 26 polymers to construct a new dataset containing 5,215,444 MMMs samples. The 8 input variables P of the new MMMs 1 He , P 1 CH4 , S 1 He / CH4 、loading、 Φ, polymer type, and MOF type were input into the XGBoost model to predict their P He and S He / CH4Considering the large size and complexity of our dataset, the 5,215,444 MMMs were divided into 26 parts, each containing 200,594 MMMs made of the same polymer. The top 6 predicted high-performance MMMs for each polymer type were screened, resulting in a total of 156 MMMs. The predicted P He and S He / CH4 The values ​​were compared with those calculated using the method for constructing the CoREMOF / polymer MMMs dataset described above. Figure 6 (a) and 9(b) show that the predicted PHe and SHe / CH4 are in good agreement with the high-throughput computer calculation results with high R2 values ​​(>0.99), demonstrating that our XGBoost model can accurately predict the helium separation performance of hMOF / polymer MMMs.

Claims

1. A method for predicting the gas separation performance of a mixed matrix membrane by machine learning, characterized in that: The steps include: Step 1, setting the physical property parameters of the mixed matrix membrane as input variables and setting the gas separation performance of the mixed matrix membrane as target variables; Step 2, using the data of the training samples to train the machine learning model to obtain a prediction model; Step 3: Use the prediction model to predict the gas separation performance of the mixed matrix membrane.

2. The method for predicting the gas separation performance of a mixed matrix membrane by machine learning according to claim 1, characterized in that: The mixed matrix membrane refers to a membrane obtained by mixing a polymer and a filler and capable of separating gases.

3. The method for predicting the gas separation performance of a mixed matrix membrane by machine learning according to claim 1, characterized in that: The filler is one of MOF material, COF material, two-dimensional sheet material, molecular sieve, carbon nanotube, graphene, metal oxide, transition metal sulfide and the like.

4. The method for predicting the gas separation performance of a mixed matrix membrane by machine learning according to claim 1, characterized in that: The polymer is selected from polysulfone (PSU), polyimide (PI), polyethersulfone (PES), polyetherimide (PEI), polycarbonate (PC), polytetrafluoroethylene (PTFE), polyvinylidenefluoride (PVDF), polyvinyl alcohol (PVA), polyacrylonitrile (PAN), polylactic acid (PLA), polystyrene (PS), polyvinyl chloride (PVC), polypropylene (PP), polyethylene terephthalate (PET), polyurethane (PU), etc.

5. The method for predicting the gas separation performance of a mixed matrix membrane by machine learning according to claim 1, characterized in that: The mass percentage of the filler in the polymer is in the range of 0.01-20%.

6. The method for predicting the gas separation performance of a mixed matrix membrane by machine learning according to claim 1, characterized in that: The gas separation refers to CO2, CH4, N2, O2, H2, He, CO, H2O, NH3, H2S, SO2, NO x , CH3OH, C3H8, C2H6, CO or a mixture of two or more gases can be separated.

7. The method for predicting the gas separation performance of a mixed matrix membrane by machine learning according to claim 1, characterized in that: The machine learning model is selected from one or a combination of decision tree (DT), random forest (RF), lightweight gradient boosting machine (LightGBM) and extreme gradient boosting (XGBoost), support vector machine (SVM), logistic regression (LogisticRegression), naive Bayes (NaiveeBayes), K-Nearest Neighbors (KNN), linear regression (LinearRegression), neural networks (NeuralNetworks), deep learning (DeepLearning), convolutional neural networks (ConvolutionalNeuralNetworks, CNN), recurrent neural networks (RecurrentNeuralNetworks, RNN), long short-term memory network (LongShort-TermMemory, LSTM), clustering algorithm (such as K-Means), principal component analysis (PCA), natural language processing model (such as BERT), generative adversarial networks (GenerativeAdversarialNetworks, GANs); decision tree (DT), random forest (RF), lightweight gradient boosting machine (LightGBM) and extreme gradient boosting (XGBoost) are preferred; extreme gradient boosting (XGBoost) is further preferred.

8. The method for predicting the gas separation performance of a mixed matrix membrane by machine learning according to claim 1, characterized in that: The physical property parameters include: polymer permeability to one component, polymer permeability to another component, polymer selectivity to two components, temperature, filler volume fraction in the mixed matrix membrane, filler global maximum pore diameter GCD Maximum cavity diameter of filler LCD Packing pore size limiting diameter PLD Packing density ρ(g / cm 3 ), filler accessible specific surface area VSA (m 2 / cm 3 ), filler accessible pore volume V p (cm 3 / g) and a combination of one or more of the filler accessible porosity Φ, the filler type, and the polymer type; preferably, a combination of one or more of the polymer permeability to one component, the polymer permeability to another component, the polymer selectivity to two components, the volume fraction of the filler in the mixed matrix membrane, the filler pore limiting diameter PLD, the filler accessible porosity Φ, the filler type, and the polymer type.

9. The method for predicting the gas separation performance of a mixed matrix membrane by machine learning according to claim 1, characterized in that: The gas separation performance refers to one or a combination of the permeability of the mixed matrix membrane to one gas, the separation selectivity of the mixed matrix membrane to two gas components; The permeability of the mixed matrix membrane to one gas is calculated by the following formula: is the permeability of the filler to gas, is the volume fraction of filler and interfacial voids, represents the permeability of the polymer, P i represents the permeability of MMM to gas; and, It is calculated by the following formula: is the permeability coefficient of gas i, is the Henry constant for gas i, is the gas i diffusion coefficient.

10. The method for predicting the gas separation performance of a mixed matrix membrane by machine learning according to claim 9, characterized in that: Henry's constant was calculated by the Widom particle insertion method under infinite dilution conditions of gas, and the diffusion coefficient was calculated from the slope of the linear region of the mean square displacement of gas molecules versus time; P represents the permeability coefficient of the gas in the filler, where i and j represent different gas molecules.

Citation Information

Cited By

  • Method for predicting micropollutant rejection rate of nanofiltration membrane and application thereof

    CN121725920A