A method for high-throughput screening of MOFs catalytic carbon dioxide cycloaddition catalysts
By using machine learning models to screen MOF catalysts, the challenge of high-throughput screening of MOF catalysts has been solved, achieving efficient and low-cost catalyst screening and data analysis, and supporting industrial applications.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING UNIV OF TECH
- Filing Date
- 2023-09-13
- Publication Date
- 2026-05-05
AI Technical Summary
Existing technologies cannot achieve high-throughput screening of MOF catalysts, especially in the field of carbon dioxide cycloaddition reactions. Traditional experimental methods are time-consuming and costly, while theoretical calculations are costly and difficult to obtain relevant performance data.
A machine learning program based on DFT calculation and low-cost GCMC simulation were used, combined with descriptors of reaction conditions, to establish a machine learning model. The TOF value was selected as the classification criterion, and a high-throughput screening model was established by selecting catalysts through machine learning algorithms.
It significantly reduces R&D time and economic costs, enables high-throughput screening of MOF catalysts, provides data analysis guidance, is applicable to different reaction conditions, and supports the transition from laboratory to industrial applications.
Smart Images

Figure CN117198415B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of chemical MOF catalysis technology, and more particularly to a method for high-throughput screening of MOF-catalyzed carbon dioxide cycloaddition catalysts. Background Technology
[0002] MOFs, short for Metal-Organic Frameworks, are a class of materials composed of metal ions and organic ligands. MOFs possess highly ordered pore structures and surface areas, characterized by tunable pore size and structure, thus finding wide applications in catalysis. Due to their structural properties, MOFs can achieve various catalytic processes with different mechanisms, including electrocatalysis, photocatalysis, and traditional thermocatalysis, such as the carbon dioxide cycloaddition reaction. The theoretically limitless possibilities and high microscopic tunability of MOFs make high-throughput screening a major challenge in their development. Other applications of MOFs, such as adsorption separation, can be simulated using low-cost Monte Carlo (GCMC) calculations [CN 115458073 A], but because thermocatalysis involves multiple factors such as heat transfer, mass transfer, and transmission, screening typically requires laboratory experimental methods. However, research on a single material typically takes several months. Since MOFs are composed of metals and organic ligands arranged in a specific topological structure, there are theoretically an infinite number of possibilities. Furthermore, the microscopic control of MOFs is complex and multifaceted. Traditional experimental methods struggle to screen and develop materials quickly, let alone achieve high-throughput screening.
[0003] While theoretical calculations, as a computational tool in chemical engineering, have been used in the development of chemical materials, the large number of atoms in the unit cell of MOF materials and the complexity of the catalytic mechanism result in high computational costs. Furthermore, theoretical calculations using density functional theory (DFT) and molecular dynamics (MD) are often insufficient to obtain data directly related to material performance. Therefore, current technologies cannot achieve high-throughput screening of MOF catalysts. To address this, we propose a method for high-throughput screening of MOF catalysts for the cycloaddition of carbon dioxide, using the catalytic reaction of carbon dioxide cycloaddition as an example. Summary of the Invention
[0004] The present invention mainly addresses the technical problems existing in the prior art and provides a method for high-throughput screening of MOFs-catalyzed carbon dioxide cycloaddition catalysts.
[0005] To achieve the above objectives, this invention employs the following technical solution: a method for high-throughput screening of MOF-catalyzed carbon dioxide cycloaddition catalysts. A machine learning model is established using a DFT-based machine learning program and low-cost GCMC simulations to obtain MOF structures and electronic properties, along with descriptors incorporating reaction conditions. The TOF value of the catalyst is selected as the classification criterion, thus establishing a high-throughput screening model. Furthermore, the application of machine learning technology provides a data analysis scheme for MOF catalyst design, further comprising the following method steps:
[0006] S1. Data preparation;
[0007] S2, Calculation of metallic charge;
[0008] S3. Calculation of specific surface area and pore volume;
[0009] S4. Machine learning algorithm selection;
[0010] S5, Algorithm Evaluation;
[0011] S6. Model data analysis.
[0012] Preferably, in S1, further:
[0013] Data sets extracted from recent papers were used as raw data for machine learning, with structural features and response condition information selected as feature descriptors. Appropriate feature descriptors reduced computational costs, enabling large-scale screening while maintaining computational accuracy.
[0014] Preferably, in S2, further:
[0015] Metal charges were calculated using the PACMOF program obtained from GitHub. The algorithm model used was the DDEC model based on the random forest algorithm provided by the program. No additional model training was performed. Before the calculation, all CIF files were processed to remove solvent and guest molecules.
[0016] Preferably, in S3, further:
[0017] Specific surface area and pore volume were calculated using 50,000 Monte Carlo calculations of giant canonical system synthesis in RASPA code. Helium was used as the probe molecule to calculate the specific surface area, and the pore volume was obtained by multiplying the helium porosity by the specific volume. The force field was selected as the UFF universal force field, and the cutoff value was set to 12.8.
[0018] Preferably, in S4, further:
[0019] As a common expression of catalytic activity, turnover frequency (TOF) was used as the target for model training. Sixteen mainstream machine learning classification models were used, including Python's SCIKIT-Learn, catboost, xgboost, and LGBM packages. This scheme used KNN, LR, DT, RF, SVM, SGD, QDA, NN, BernoulliNB (BNB), MultinomialNB (MNB), CatBoost (Cat), LightGBM (LGBM), XGBoost (XGB), GBDT, ET, and AdaBoost (Ada) models. All model training was divided into training and test sets at an 80 / 20 ratio, with a random seed set to 1. Grid search hyperparameters were used to perform triple cross-validation on the training set.
[0020] Preferably, in step S5, further:
[0021] The performance of each machine learning algorithm was evaluated by calculating the model's accuracy, precision, recall, and F1 score, resulting in a machine learning model with an accuracy as high as 97%, such as... Figure 2 As shown.
[0022] Preferably, during the high-throughput verification of the computational model, a virtual MOF structure was constructed using Topological, Based, Crystal, and Constructor (ToBaCCo3.0) code, along with 100 actual MOF structures obtained from the University of Cambridge database.
[0023] Preferably, in S6, further:
[0024] Data analysis of machine learning algorithms was performed using SHapley Additive exPlanations and Partial Dependence Plot, yielding empirical guidance with chemical significance.
[0025] Beneficial effects
[0026] This invention provides a method for high-throughput screening of MOF-catalyzed carbon dioxide cycloaddition catalysts. It has the following beneficial effects:
[0027] (1) Compared with the prior art, this invention provides a method for high-throughput screening of MOFs catalysts for carbon dioxide cycloaddition, which solves the problem of high-throughput catalyst screening technology in the field of MOFs catalysis, especially in the field of carbon dioxide cycloaddition.
[0028] (2) This method establishes a complete workflow for data extraction, data preparation, metal charge calculation, machine learning algorithm selection, algorithm evaluation, and model data analysis, which significantly reduces R&D time and economic costs compared with experimental research.
[0029] (3) This method provides a model data analysis method, and the conclusions drawn can effectively guide the development, synthesis and improvement of MOF catalysts.
[0030] (4) This method evaluated different machine learning models and invoked a variety of feature descriptors with practical chemical significance. The established method has guiding and general applicability for high-throughput screening of other MOF catalysts.
[0031] (5) Since the reaction conditions are also the descriptors of the model, it is feasible to screen catalysts under different temperature, pressure and co-catalyst conditions. This process effectively realizes the transition from laboratory to industrial application. Attached Figure Description
[0032] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings in the following description are merely exemplary, and those skilled in the art can derive other embodiments based on the provided drawings without creative effort.
[0033] The structures, proportions, sizes, etc. illustrated in this specification are only for the purpose of assisting those skilled in the art in understanding and reading the content disclosed herein, and are not intended to limit the conditions under which the present invention can be implemented. Therefore, they have no substantial technical significance. Any modifications to the structure, changes in the proportions, or adjustments to the size, without affecting the effects and objectives that the present invention can produce, should still fall within the scope of the technical content disclosed in the present invention.
[0034] Figure 1 This is the high-throughput screening model of the present invention;
[0035] Figure 2 This describes the data distribution of the dataset used in this invention.
[0036] Figure 3 This invention relates to the machine learning model algorithm and model evaluation;
[0037] Figure 4 This is a representative structural diagram of the highly active MOF of the present invention;
[0038] Figure 5This is the analytical model established by the SHapley Additive exPlanations and Partial Dependence Plot in this invention. Detailed Implementation
[0039] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0040] Example 1: A method for high-throughput screening of MOF-catalyzed carbon dioxide cycloaddition catalysts, the design concept of which is as follows: Figure 1 As shown, a machine learning model was established using a DFT-based machine learning program and low-cost GCMC simulations to obtain the structure and electronic properties of MOFs, along with descriptors related to reaction conditions. The TOF value of the catalyst was selected as the classification criterion, and a high-throughput screening model was established. Furthermore, the application of machine learning technology also provides a data analysis scheme for MOF catalyst design, further including the following methodological steps:
[0041] S1. Data preparation;
[0042] S2, Calculation of metallic charge;
[0043] S3. Calculation of specific surface area and pore volume;
[0044] S4. Machine learning algorithm selection;
[0045] S5, Algorithm Evaluation;
[0046] S6. Model data analysis.
[0047] In S1, further:
[0048] The dataset extracted from 71 recent papers was used as the raw data for machine learning, and structural features and reaction condition information were selected as feature descriptors. As shown in Table 1, it specifically includes the metal species, ligand characteristics, reaction temperature, reaction pressure, type and amount of co-catalyst, substrate species, turnover frequency (TOF), and MOF structure files for further calculation of metal charge, specific surface area, and pore volume.
[0049] Table 1. Explanation and meaning of descriptors and targets.
[0050]
[0051]
[0052] In S2, further:
[0053] Metal charges are calculated by passing the structure files of MOFs to the PACMOF program obtained from Github. The algorithm model is the DDEC model based on the random forest algorithm provided by the program. No additional model training is performed. Before the calculation, all MOF structure files are processed to remove solvent and guest molecules.
[0054] In S3, further:
[0055] Specific surface area and pore volume were calculated by performing GCMC calculations on MOF structure files for 50,000 cycles of the giant canonical system using RASPA code. Helium was used as the probe molecule to calculate the specific surface area, and the pore volume was obtained by multiplying the helium porosity by the specific volume. The force field was selected as the UFF universal force field, and the cutoff value was set to 12.8.
[0056] The data distribution obtained from S1-S3 above is as follows: Figure 2 As shown.
[0057] In S4, further:
[0058] As a common expression of catalytic activity, turnover frequency (TOF) was used as the target for model training. Sixteen mainstream machine learning classification models were used, including Python's SCIKIT-Learn, catboost, xgboost, and LGBM packages. This scheme used KNN, LR, DT, RF, SVM, SGD, QDA, NN, BernoulliNB (BNB), MultinomialNB (MNB), CatBoost (Cat), LightGBM (LGBM), XGBoost (XGB), GBDT, ET, and AdaBoost (Ada) models. All model training was divided into training and test sets at an 80 / 20 ratio, with a random seed set to 1. Grid search hyperparameters were used to perform triple cross-validation on the training set.
[0059] In S5, further:
[0060] The performance of each machine learning algorithm was evaluated by calculating the model's accuracy, precision, recall, and F1 score, and the Random Forest machine learning model with an accuracy as high as 97% was selected. For example... Figure 3As shown. During high-throughput validation of the computational model, the Topological Based Crystal Constructor (ToBaCCo3.0) code was used to construct, but is not limited to, 12415 virtual MOF structures and 100 actual MOF structures. Finally, within one day, 237 (virtual) + 2 (actual) MOF catalysts with potential catalytic activity were screened. Some of their MOF structures are shown below. Figure 4 As shown.
[0061] In S6, further:
[0062] Data analysis of the machine learning algorithm was performed using Shapley Additive exPlanations and Partial Dependence Plot. The results showed that the presence of Y element generally increases reactivity, and that small molecule substrates such as propylene oxide typically exhibit better TOF values compared to large molecule substrates such as styrene oxide. Specific data are shown below. Figure 5 As shown.
[0063] Example 2: A method for high-throughput screening of MOFs-catalyzed carbon dioxide cycloaddition catalysts, based on Example 1, wherein in step S1, data can be acquired through self-experimentation, database collection, etc.
[0064] In S2, the charge calculation can use the nbo charge, LUMO-HOMO energy level and other energy representations of interatomic interactions to replace the PAC charge;
[0065] In S3, the specific surface area and pore volume can be obtained through database acquisition, theoretical calculation, experiments, etc.
[0066] In S4, the machine learning algorithm can be calculated using various other algorithms, such as ANN.
[0067] In step S5, the MOFs materials used for screening include, but are not limited to, experimentally obtained materials, the CoREMOF database, and other MOFs structure generation algorithms.
[0068] In S6, the analysis of the machine learning model includes, but is not limited to, SHapley Additive exPlanations and Partial Dependence Plot.
[0069] Working principle of the invention:
[0070] This technique, based on datasets extracted from the literature, utilizes a machine learning program based on DFT calculations and low-cost GCMC simulations to obtain MOF structures and electronic properties, along with descriptors combining reaction conditions, to select catalyst TOF values as a classification criterion, establishing a machine learning model for high-throughput screening. The core of this process lies in the comprehensive description of the chemical processes involved in catalysis and the inherent characteristics of the MOF materials themselves through the selection of reasonable, low-cost descriptors, achieving high-accuracy machine modeling and screening. Furthermore, all features are selected from readily available ones to ensure high-throughput screening. Finally, the machine learning analysis model is used to guide chemical synthesis and MOF catalyst design by interpreting the algorithm's operational logic. In summary, this study leverages existing chemical experience and low-cost computation, selects feature descriptors, and employs machine learning techniques to build a model, achieving high-throughput screening and data analysis of MOF materials.
[0071] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.
Claims
1. A method for high-throughput screening of MOF-catalyzed carbon dioxide cycloaddition catalysts, characterized in that: A machine learning model was established using a DFT-based machine learning program and low-cost GCMC simulations to obtain the structure and electronic properties of MOFs, along with descriptors related to reaction conditions. The TOF value of the catalyst was selected as the classification criterion, and a high-throughput screening model was established. Furthermore, the application of machine learning technology also provides a data analysis scheme for MOF catalyst design, further including the following methodological steps: S1. Data preparation; S2, Calculation of metallic charge; S3. Calculation of specific surface area and pore volume; S4. Machine learning algorithm selection; S5, Algorithm Evaluation; S6. Model data analysis; In S2, further: Metal charges were calculated using the PACMOF program obtained from Github. The algorithm model used was the DDEC model based on the random forest algorithm provided by the program. No additional model training was performed. Before the calculation, all MOF structure files were preprocessed to remove solvent and guest molecules from the structure. In S3, further: Specific surface area and pore volume were calculated using Monte Carlo calculations with 50,000 cycles of the giant canonical system in RASPA code. Helium was used as the probe molecule to calculate the specific surface area, and the pore volume was obtained by multiplying the helium porosity by the specific volume. The force field was selected as the UFF universal force field, and the cutoff value was set to 12.
8.
2. The method for high-throughput screening of MOF-catalyzed carbon dioxide cycloaddition catalysts according to claim 1, characterized in that: In S1, further: The dataset extracted from the paper was used as the raw data for machine learning. Structural features and reaction condition information were selected as feature descriptors. The atomic partial charge of the metal center was calculated using the machine learning model PACMOF based on density functional theory (DFT) to describe the Lewis acidity of the metal open sites. The specific surface area and porosity of the structure were calculated using RASPA Monte Carlo (GCMC) simulation to describe the three-dimensional properties of MOF as feature inputs.
3. The method for high-throughput screening of MOF-catalyzed carbon dioxide cycloaddition catalysts according to claim 1, characterized in that: In S4, further: As a common expression of catalytic activity, turnover frequency (TOF) was used as the target for model training. Sixteen mainstream machine learning classification models were used, employing Python's SCIKIT-Learn, catboost, xgboost, and LGBM packages. Specifically, this approach used KNN, LR, DT, RF, SVM, SGD, QDA, NN, BernoulliNB (BNB), MultinomialNB (MNB), CatBoost (Cat), LightGBM (LGBM), XGBoost (XGB), GBDT, ET, and AdaBoost (Ada) models. All model training was divided into training and test sets at an 80 / 20 ratio, with a random seed of 1. Hyperparameters were searched using a grid search, and triple cross-validation was performed on the training set.
4. The method for high-throughput screening of MOF-catalyzed carbon dioxide cycloaddition catalysts according to claim 1, characterized in that: In S5, further: The performance of each machine learning algorithm is evaluated by calculating the accuracy, precision, recall, and F1 score of the model, resulting in a preferred machine learning model with an accuracy of up to 97%.
5. The method for high-throughput screening of MOF-catalyzed carbon dioxide cycloaddition catalysts according to claim 4, characterized in that: During the high-throughput verification of the computational model, the code of Topological Based CrystalConstructor (ToBaCCo3.0) was used to construct, but was not limited to, 12,415 virtual MOF structures and 100 actual MOF structures obtained from the Cambridge University database. The model successfully screened out MOF materials with excellent catalytic activity.
6. The method for high-throughput screening of MOF-catalyzed carbon dioxide cycloaddition catalysts according to claim 1, characterized in that: In step S6, the machine learning algorithm is further analyzed using SHapley Additive exPlanations and PartialDependence Plot to derive empirical guidance with chemical significance.
Citation Information
Patent Citations
Machine learning method for hydrophobic MOF adsorbing DMMP
CN115458073A
Machine learning algorithm optimization method based on MOF, storage medium and computing device
CN114898826A
Method for high-throughput screening of metal organic framework membrane based on machine learning assistance
CN115620835A