A method and device for screening adsorption materials based on machine learning

Through machine learning-based adsorbent material screening methods, the problem of difficulty in quickly screening the best adsorbent materials in water treatment is solved, efficient and intelligent material design and screening is achieved, research costs and cycles are reduced, and the synthesis of adsorbent materials is guided.

CN113963754BActive Publication Date: 2025-05-23NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111073322.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-14
Publication Date
2025-05-23
Estimated Expiration
2041-09-14

AI Technical Summary

Technical Problem

The prior art is difficult to quickly and efficiently screen out the best adsorbent materials in water treatment, especially when facing a variety of complex organic pollutants, traditional methods require long research cycles and high experimental costs.

Method used

Using machine learning-based adsorption material screening method, the original data set is established and pre-processed by obtaining descriptors of materials, pollutants and adsorption process. Then, multiple machine learning models are used for training to determine the optimal hyperparameters and sampling methods, identify the key parameters that affect adsorption performance, and finally screen out the best adsorption material.

Benefits of technology

It realizes intelligent design and rapid screening of the best adsorbent materials in the field of water treatment, reduces research cycle and cost, improves the interpretability of adsorption effects, and guides the synthesis of specific adsorbent materials.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113963754B_ABST
    Figure CN113963754B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for screening adsorbent materials based on machine learning, including obtaining material descriptors, pollutant descriptors and adsorption process descriptors, establishing original data sets respectively, and preprocessing the original data; inputting the preprocessed original data into multiple machine learning models for training, and determining the optimal hyperparameters according to the training results; evaluating the performance of the machine learning model, and selecting the best prediction model; identifying the key parameters affecting the adsorption performance through feature engineering; establishing a candidate material library, inputting the pollutants to be tested into the best prediction model, and locating the material with the best adsorption performance; inputting the best material into the candidate material library, and obtaining the synthesis method of the best material. The present invention can quickly and accurately locate a certain material by screening the best material, and return to the CCDC database to obtain its relevant information and related literature of the first synthesis, so as to guide the material synthesis, and then realize the intelligent design of the adsorbent material.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and device for screening adsorbent materials, and in particular to a method and device for screening adsorbent materials based on machine learning. Background Art

[0002] Adsorption has been widely studied in the field of water treatment due to its high efficiency and convenience. Over the years, various types of adsorbents have emerged, and the number of various types of adsorbents such as biochar, molecular sieves, and metal-organic frameworks has shown explosive growth.

[0003] There are many adsorbent materials, and efficient adsorbents for different pollutants are constantly being discovered. However, the efficient adsorbents discovered by experiments today are only the tip of the iceberg of all adsorbent materials that may exist in reality. At the same time, there are many types of pollutants in sewage, and it is difficult for a single adsorbent to guarantee the same high adsorption efficiency for different pollutants in sewage. If the adsorption performance of all existing materials for different pollutants is verified through traditional adsorption experiments, it means that a long research cycle and high experimental costs are required. Therefore, through traditional experimental methods, selecting the best adsorbent for a specific pollutant is like looking for a needle in a haystack, and repeated trial and error.

[0004] Therefore, it is necessary to establish certain standards to provide the possibility for efficient screening of adsorption materials. At present, the Materials Genome Initiative (MGI) is leading a new model of material research and development. One of the important challenges is to integrate high-throughput computing technology and calculate material structure information based on the concept of materials genomics to identify the best possible materials. Metal organic frameworks (MOFs) are modular porous materials that can be freely spliced ​​by metal centers and organic ligands to form a periodic network structure. In theory, they can be assembled into an almost unlimited number of materials with the characteristics of diversity, regularity, and designability. At the same time, the descriptors of MOF structural information can be used to locate a specific material. Taking advantage of this advantage, researchers have analyzed the structural information of known MOFs synthesized experimentally, and at the same time, with the help of chemical theory and computer technology, they have constructed thousands of MOFs that may be synthesized, and established databases for screening, such as the coREMOFs and hMOFs databases established by Northwestern University, which lay the foundation for high-throughput computational screening methods based on molecular simulation and have been widely used in the fields of gas adsorption, separation and storage.

[0005] Although the above technologies provide the possibility for rapid screening of materials, on the one hand, the technology is limited to gas adsorption and cannot be extended to the removal of complex organic matter in water, and there is a research gap in the field of water treatment; on the other hand, with the continuous growth of the number of MOFs, high-throughput screening based on molecular simulation is often limited by the huge MOFs database and limited computing resources, and the support of big data computing plays a pivotal role. Therefore, machine learning provides a strong guarantee for the screening of MOF materials for the removal of pollutants in water, and rapid material screening technology based on machine learning is urgently needed to be developed. Summary of the invention

[0006] Purpose of the invention: The purpose of the present invention is to provide a fast, accurate and efficient method for screening adsorbent materials based on machine learning; another purpose of the present invention is to provide a device for screening adsorbent materials based on machine learning.

[0007] Technical solution: The adsorption material screening method based on machine learning of the present invention comprises the following steps:

[0008] (1) Obtain material descriptors, pollutant descriptors, and adsorption process descriptors, establish original data sets respectively, and preprocess the original data;

[0009] (2) inputting the preprocessed raw data into multiple machine learning models for training, and determining the optimal hyperparameters based on the training results; evaluating the performance of the machine learning models and selecting the best prediction model; evaluating point sampling and group sampling and selecting the optimal sampling method, wherein the group sampling is to group the raw data in units of adsorption isotherms, and the point sampling is to group the raw data without grouping; identifying the key parameters affecting the adsorption performance through feature engineering;

[0010] (3) establishing a candidate material library, the candidate material library including the material descriptors; inputting the pollutant to be tested into the optimal prediction model to locate the material with the best adsorption performance; inputting the best material into the candidate material library to obtain a synthesis method for the best material.

[0011] Furthermore, in step (1), the adsorbent material information contains a CIF structure file; if the adsorbent material information does not contain a CIF structure file, Material studio is used to process the adsorbent material information to obtain a CIF structure file. The adsorbent material that does not contain a CIF structure file is a modified adsorbent material. The CIF structure file of the raw material before modification can be found in the structure database. Material studio is used to edit and modify the CIF structure file of the raw material to simulate the modified adsorbent material so that the information described in the CIF structure file restores the characteristics of the modified material in the literature as much as possible.

[0012] Furthermore, in step (1), the adsorption material descriptor includes a mass specific surface area descriptor, a volume specific surface area descriptor, a topological structure descriptor, a metal center descriptor and an organic ligand descriptor.

[0013] Furthermore, in step (1), the pollutant descriptors include a molecular molar refractive index descriptor, a molecular dipole / polarizability descriptor, a hydrogen bonding proton acceptor ability descriptor, a hydrogen bonding proton donor ability descriptor and a molecular volume descriptor.

[0014] Furthermore, in step (1), the adsorption process descriptor is an adsorption amount descriptor in the adsorption isotherm.

[0015] The specific method for obtaining the adsorption material description, pollutant descriptor and adsorption process descriptor is as follows: a large number of existing research articles are obtained, various information related to the materials used, the pollutants studied, the adsorption conditions and the adsorption performance are obtained from the articles, and the collected data are calculated and information extracted. According to the description of the material in the article, the CIF structure file is obtained from the Cambridge University Structure Database. If the research article simply modifies the original material, the Material studio modifies the original structure file, and then uses the structure file to calculate its mass specific surface area, volume specific surface area, topological structure and other physical descriptors and metal center, organic ligand chemical descriptors as characteristic descriptors of the material; according to the description of pollutants in the article, the Abraham descriptors E, S, A, B, and V of the pollutants are obtained from the compound database as descriptors of the pollutant characteristics. The above five descriptors respectively characterize the molecular molar refractive index, molecular dipole / polarizability, hydrogen bond proton acceptor capacity, hydrogen bond proton donor capacity, and molecular volume of the pollutants; according to the description of the adsorption experiment in the article, the adsorption conditions such as temperature, pH, pollutant concentration, system volume, and solid-liquid ratio are crawled, and GetData is used to obtain the data points in the adsorption isotherm as descriptors of the adsorption process.

[0016] Furthermore, the data preprocessing method in step (1) includes data cleaning, rearrangement, and normalization, wherein data cleaning is used to improve the quality of the original data set, statistically analyze the distribution of the original data set, check the data structure, delete abnormal values ​​and outliers, and use the DropNA function to delete entries with missing values; use the shuffle function to rearrange the original data, use the train_test_split function to split the training set and the test set according to different ratios, determine the best data splitting ratio, and select the best validation set sampling method;

[0017] Normalize the training feature data and normalize the test set data according to the same rules. On the one hand, it speeds up the training processing and calculation process, and on the other hand, it prevents the deviation that may be caused by data of different dimensions. The normalization formula is as follows:

[0018]

[0019] Lmax, Lmin, x, xmin, xmax, and xstd are the upper and lower bounds of standardization (usually 1 and -1), the input value, the minimum value of x, the maximum value of x, and the standardized x, respectively.

[0020] Furthermore, the machine learning model in step (1) includes feedforward neural network, random forest, gradient boosting tree, and multi-granularity cascade forest.

[0021] Furthermore, in step (2), the method for determining the optimal hyperparameters is Gridsearch;

[0022] Furthermore, the performance of different models was evaluated and the best prediction model was selected. 2 The best prediction model was evaluated by R and RMSE. 2 The calculation formula of RMSE is as follows:

[0023]

[0024] Where n represents the total number of data sets, and yi are the best model prediction value and true value of the ith data, respectively. All predicted values The average value of .

[0025] Furthermore, in step (2), the Shapley value in feature engineering is used as an indicator to evaluate the contribution of feature variables, and the key parameters affecting the adsorption performance are identified through the retention of the model prediction performance; specifically, the Shapley values ​​of different features of the model are calculated using the SHAP package to evaluate the contribution of different features to the prediction results, while some features are eliminated to evaluate the retention effect of the model prediction performance, identify the key parameters affecting the adsorption performance, and perform feature compression.

[0026] Furthermore, in step (3), if there is no corresponding optimal material synthesis method in the candidate material library, the optimal material synthesis method is obtained by obtaining characteristic information of the metal center and organic ligand of the optimal material.

[0027] Furthermore, the construction of the candidate material library includes using the coREMOFs database or hMOFs database of Northwestern University, downloading CIF literature and json files, parsing related files, calculating and obtaining material descriptors, and establishing a candidate material library for screening potential adsorption materials.

[0028] The method of using the candidate material library includes selecting the pollutant to be removed, using the German UFZ-LSER database to obtain its Abraham descriptor, inputting the description information of the pollutant into the established model for prediction, and obtaining a series of removal effect data; statistically analyzing the removal effects of all materials to locate the material with the best removal performance; the material is a known material or a potential material, and the known material can return to the corresponding related material in the candidate material library based on the material information, and return to the CCDC database or research literature to find the synthesis method; the potential material is a material that exists in theory but has not been synthesized, and the potential material guides the material synthesis by obtaining the characteristic information of the metal center and organic ligand of the potential material.

[0029] The device using the above-mentioned machine learning-based adsorption material screening method includes a data set construction module, which is used to obtain descriptors of materials, pollutants and adsorption processes, perform data cleaning, and establish an original data set; a model pre-training module, which is used to generate multiple initial models containing hyperparameters according to different algorithms, and adjust the parameters to select the optimal hyperparameters; a model construction module, which is used to select the best prediction model, the most suitable training set and test set division ratio, and the optimal sampling method; a feature engineering module, which is used to compress features and identify key parameters that affect model performance and adsorption effect; a testing module, which is used to input a test set to verify the predictive performance of the model; a material screening module, which is used to screen the best adsorption material for specific pollutants; and a guided synthesis module, which is used to input the best adsorption material and output the synthesis method of the best material.

[0030] Furthermore, the synthesis guidance module includes a candidate material database module.

[0031] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages:

[0032] (1) In the field of water treatment, realize the intelligent design of efficient MOFs materials and quickly screen the best water treatment adsorption materials. In the past, existing research on the intelligent design of MOFs materials focused on gas adsorption, classification and energy storage, but there were deficiencies in the field of water phase adsorption;

[0033] (2) When faced with untested adsorbent materials, the optimal adsorption conditions can be quickly located and precisely controlled. In the process of model building, the parameters related to the adsorption conditions are used as feature inputs for model training, which can open up broad prospects for the rapid application of unknown adsorbent materials.

[0034] (3) To improve the interpretability of the adsorption effect, the adsorption mechanism can be explained from the perspective of structure-activity relationship. With the help of chemical calculations, the structural parameters of the material can be analyzed using the structure files of MOFs to explore the relationship between the material structural characteristics and the adsorption effect.

[0035] (4) Efficiently guide the synthesis of specific adsorbent materials. The existing technology can only predict the adsorption effect of materials, but cannot quickly screen potential high-performance adsorbent materials through the information of specified pollutants. The present invention uses the concept of the Material Genome Project to introduce the descriptors of materials into the model. The best material obtained by screening can quickly and accurately locate a certain material, return to the CCDC database to obtain its relevant information and related literature on the first synthesis, so as to guide material synthesis and realize the intelligent design of adsorbent materials. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a schematic diagram of the process of the present invention;

[0037] Figure 2 is the prediction performance of the machine learning model of the present invention, expressed as the degree of fit between the actual value and the predicted value;

[0038] Figure 3 is the Shapley value of each feature of the present invention;

[0039] Figure 4 It is the probability distribution histogram of the adsorption performance of the candidate materials in Example 2. DETAILED DESCRIPTION

[0040] The technical solution of the present invention is further described below in conjunction with the accompanying drawings.

[0041] Example 1

[0042] Screening of MOF materials for the removal of non-steroidal anti-inflammatory drugs - taking naproxen as an example. Figure 1 As shown, the specific steps are as follows:

[0043] (1) Training database establishment and data preprocessing

[0044] Training database establishment: crawl MOFs material information in existing research articles, download the corresponding material structure files in the CCDC database, calculate the material's physical descriptors (mass specific surface area, volume specific surface area, topological structure, etc.) and chemical descriptors (metal center, organic ligand, etc.) as the material's characteristic descriptors; obtain the description of pollutants in the article, and obtain the Abraham descriptors (E, S, A, B, V) of pollutants in the compound database as descriptors of pollutant characteristics; obtain the pH, temperature, volume, solid-liquid ratio and other adsorption conditions in the article, and use GetData to crawl the concentration and adsorption amount data in the adsorption isotherm of the article, calculate the removal rate, and use it as the adsorption process parameter. Combine the above three parts of data to establish a database for training the model.

[0045] Data cleaning and data splitting: Statistically analyze the distribution of training data, perform data cleaning on the original data, and delete data entries with missing values ​​and outliers; select the best split ratio of training set and test set, and select a split ratio of 7:3; normalize the training set, and normalize the test set with the same rules.

[0046] (2) Determination of the optimal prediction model

[0047] Model construction and optimal hyperparameter determination: The training data was substituted into a variety of machine learning models including feedforward neural network, random forest, gradient boosting tree and multi-granularity cascade forest for training, and the optimal hyperparameters were selected using Gridsearch. 2 , RMSE and other evaluation indicators to compare the prediction performance of different models, and select the gradient boosting tree as the most suitable model, such as Figure 2 The fit between the predicted value and the actual value is 0.982.

[0048] Determination of the optimal sampling method: Explore the effect of the group selection sampling method on the prediction performance. Use the adsorption isotherm as the unit, group the original training data, split the training set and the test set, and establish a model to predict unknown samples. Comparing the prediction performance of the sampling methods of point selection and group selection, it shows that group selection has better prediction performance for the test set, and group selection is determined as the sampling method of the model.

[0049] Feature Engineering: Use Python's SHAP package to calculate the Shapley values ​​of different features to characterize the contribution of different features to the model prediction results. The contribution of each feature is as follows: Figure 3 As shown, concentration contributes the most to the prediction results, followed by the mass specific surface area and volume specific surface area of ​​the material, and then the molecular volume of the pollutant, hydrogen bond proton acceptor capacity, molecular molar refractive index, hydrogen bond proton donor capacity and molecular dipole / polarizability. At the same time, some features were deleted, and the key parameters were identified through the retention of the model prediction performance. It was found that the prediction accuracy of the model that deleted the concentration feature decreased most significantly, decreasing to 55.7% of the original; in addition, the model that included materials, pollutant characteristics, and adsorption process parameters simultaneously obtained the best performance retention. In summary, concentration, specific surface area of ​​materials, molecular volume of pollutants, and hydrogen bond proton acceptor capacity are the key factors affecting the removal effect.

[0050] (3) Screening of optimal materials for naproxen removal

[0051] 800 material structure files were selected from the coREMOFs database, and their structural parameters were obtained through calculation and analysis. The Abraham descriptor of naproxen was queried through the German UFZ-LSER database, and certain adsorption conditions were given and input into the model to obtain the matrix data of the removal effects of all materials. The MOFs material with the best removal effect was located, and the relevant information and related literature on the initial synthesis were returned to the CCDC database based on its characteristic data. The material was obtained by repeated synthesis and adsorption experimental verification.

[0052] Example 2

[0053] Screening of MOF materials for removing antibiotics - taking tetracycline as an example. The specific steps are as follows:

[0054] (1) Training database establishment and data preprocessing

[0055] Training database establishment: crawl MOFs material information in existing research articles, download the corresponding material structure files in the CCDC database, calculate the material's physical descriptors (mass specific surface area, volume specific surface area, topological structure, etc.) and chemical descriptors (metal center, organic ligand, etc.) as the material's characteristic descriptors; obtain the description of pollutants in the article, and obtain the Abraham descriptors (E, S, A, B, V) of pollutants in the compound database as descriptors of pollutant characteristics; obtain the pH, temperature, volume, solid-liquid ratio and other adsorption conditions in the article, and use GetData to crawl the concentration and adsorption amount data in the adsorption isotherm of the article, calculate the removal rate, and use it as the adsorption process parameter. Combine the above three parts of data to establish a database for training the model.

[0056] Data cleaning and data splitting: Statistically analyze the distribution of training data, perform data cleaning on the original data, and delete data entries with missing values ​​and outliers; select the best split ratio of training set and test set, and select a split ratio of 7:3; normalize the training set, and normalize the test set with the same rules.

[0057] (2) Determination of the optimal prediction model

[0058] Model construction and optimal hyperparameter determination: The training data was substituted into a variety of machine learning models including feedforward neural network, random forest, gradient boosting tree and multi-granularity cascade forest for training, and the optimal hyperparameters were selected using Gridsearch. 2 , RMSE and other evaluation indicators to compare the prediction performance of different models, and select the gradient boosting tree as the most suitable model, such as Figure 2 The fit between the predicted value and the actual value is 0.982.

[0059] Determination of the optimal sampling method: Explore the effect of the group selection sampling method on the prediction performance. Use the adsorption isotherm as the unit, group the original training data, split the training set and the test set, and establish a model to predict unknown samples. Comparing the prediction performance of the sampling methods of point selection and group selection, it shows that group selection has better prediction performance for the test set, and group selection is determined as the sampling method of the model.

[0060] Feature Engineering: Use Python's SHAP package to calculate the Shapley values ​​of different features to characterize the contribution of different features to the model prediction results. The contribution of each feature is as follows: Figure 3 As shown, concentration contributes the most to the prediction results, followed by the mass specific surface area and volume specific surface area of ​​the material, and then the molecular volume of the pollutant, hydrogen bond proton acceptor capacity, molecular molar refractive index, hydrogen bond proton donor capacity and molecular dipole / polarizability. At the same time, some features were deleted, and the key parameters were identified through the retention of the model prediction performance. It was found that the prediction accuracy of the model that deleted the concentration feature decreased most significantly, decreasing to 55.7% of the original; in addition, the model that included materials, pollutant characteristics, and adsorption process parameters simultaneously obtained the best performance retention. In summary, concentration, specific surface area of ​​materials, molecular volume of pollutants, and hydrogen bond proton acceptor capacity are the key factors affecting the removal effect.

[0061] (3) Screening of optimal materials for tetracycline removal

[0062] 800 material structure files were selected from the coREMOFs database, and their structural parameters were obtained by calculation and analysis. The Abraham descriptor of tetracycline was queried through the German UFZ-LSER database, and certain adsorption conditions were given and input into the model to obtain the matrix data of the removal effect of all materials. The distribution of adsorption effect is shown in Figure 4 As shown, the MOFs material with the best removal effect is located, and according to its characteristic data, the CCDC database is returned to search for relevant information and related literature on initial synthesis. The material is repeatedly synthesized and verified by adsorption experiments.

[0063] Example 3

[0064] Screening of MOF materials for removing endocrine disruptors - taking nonylphenol as an example. The specific steps are as follows:

[0065] (1) Training database establishment and data preprocessing

[0066] Training database establishment: crawl MOFs material information in existing research articles, download the corresponding material structure files in the CCDC database, calculate the material's physical descriptors (mass specific surface area, volume specific surface area, topological structure, etc.) and chemical descriptors (metal center, organic ligand, etc.) as the material's characteristic descriptors; obtain the description of pollutants in the article, and obtain the Abraham descriptors (E, S, A, B, V) of pollutants in the compound database as descriptors of pollutant characteristics; obtain the pH, temperature, volume, solid-liquid ratio and other adsorption conditions in the article, and use GetData to crawl the concentration and adsorption amount data in the adsorption isotherm of the article, calculate the removal rate, and use it as the adsorption process parameter. Combine the above three parts of data to establish a database for training the model.

[0067] Data cleaning and data splitting: Statistically analyze the distribution of training data, perform data cleaning on the original data, and delete data entries with missing values ​​and outliers; select the best split ratio of training set and test set, and select a split ratio of 7:3; normalize the training set, and normalize the test set with the same rules.

[0068] (2) Determination of the optimal prediction model

[0069] Model construction and optimal hyperparameter determination: The training data was substituted into a variety of machine learning models including feedforward neural network, random forest, gradient boosting tree and multi-granularity cascade forest for training, and the optimal hyperparameters were selected using Gridsearch. 2 , RMSE and other evaluation indicators to compare the prediction performance of different models, and select the gradient boosting tree as the most suitable model, such as Figure 2 The fit between the predicted value and the actual value is 0.982.

[0070] Determination of the optimal sampling method: Explore the effect of the group selection sampling method on the prediction performance. Use the adsorption isotherm as the unit, group the original training data, split the training set and the test set, and establish a model to predict unknown samples. Comparing the prediction performance of the sampling methods of point selection and group selection, it shows that group selection has better prediction performance for the test set, and group selection is determined as the sampling method of the model.

[0071] Feature Engineering: Use Python's SHAP package to calculate the Shapley values ​​of different features to characterize the contribution of different features to the model prediction results. The contribution of each feature is as follows: Figure 3As shown, concentration contributes the most to the prediction results, followed by the mass specific surface area and volume specific surface area of ​​the material, and then the molecular volume of the pollutant, hydrogen bond proton acceptor capacity, molecular molar refractive index, hydrogen bond proton donor capacity and molecular dipole / polarizability. At the same time, some features were deleted, and the key parameters were identified through the retention of the model prediction performance. It was found that the prediction accuracy of the model that deleted the concentration feature decreased most significantly, decreasing to 55.7% of the original; in addition, the model that included materials, pollutant characteristics, and adsorption process parameters simultaneously obtained the best performance retention. In summary, concentration, specific surface area of ​​materials, molecular volume of pollutants, and hydrogen bond proton acceptor capacity are the key factors affecting the removal effect.

[0072] (3) Screening of optimal materials for nonylphenol removal

[0073] 800 material structure files were selected from the coREMOFs database, and their structural parameters were obtained through calculation and analysis. The Abraham descriptor of nonylphenol was queried through the German UFZ-LSER database, and certain adsorption conditions were given and input into the model to obtain the matrix data of the removal effects of all materials. The MOFs material with the best removal effect was located, and the relevant information and related literature on the initial synthesis were returned to the CCDC database based on its characteristic data. The material was obtained by repeated synthesis and adsorption experimental verification.

[0074] So far, the technical scheme of the present invention has been schematically described in combination with the preferred embodiments of three types of pollutants, namely, non-steroidal anti-inflammatory drugs, antibiotics and endocrine disruptors, but the protection scope of the present invention is obviously not limited to the specific embodiments of these three types of pollutants. Without departing from the spirit or principle of the present invention, as long as the pollutants to be removed have Abraham descriptors, they are within the scope that can be calculated by the present invention, and those skilled in the art can make equivalent changes, substitutions or improvements to the relevant technical features. Without departing from the purpose of the present invention, the technical schemes after these changes, substitutions or improvements will fall within the protection scope of the present invention.

Claims

1. A method for screening adsorbent materials based on machine learning, It is characterized in that The following steps are involved: (1) Obtain material descriptors, pollutant descriptors, and adsorption process descriptors, establish original data sets respectively, and preprocess the original data; Acquire a material descriptor from adsorption material information, wherein the adsorption material information includes a CIF structure file; If the adsorption material information does not contain a CIF structure file, the adsorption material information is processed using Material Studio to obtain a CIF structure file; the adsorption material descriptor includes a mass specific surface area descriptor, a volume specific surface area descriptor, a topological structure descriptor, a metal center descriptor and an organic ligand descriptor; The pollutant descriptors include a molecular molar refractive index descriptor, a molecular dipole / polarizability descriptor, a hydrogen bonding proton acceptor ability descriptor, a hydrogen bonding proton donor ability descriptor, and a molecular volume descriptor; The adsorption process descriptor is the adsorption amount descriptor in the adsorption isotherm; (2) inputting the preprocessed raw data into multiple machine learning models for training, and determining the optimal hyperparameters based on the training results; evaluating the performance of the machine learning models and selecting the best prediction model; evaluating point sampling and group sampling and selecting the optimal sampling method, wherein the group sampling is to group the raw data in units of adsorption isotherms, and the point sampling is to group the raw data without grouping; identifying the key parameters affecting the adsorption performance through feature engineering; (3) establishing a candidate material library, the candidate material library including the material descriptors; inputting the pollutant to be tested into the optimal prediction model to locate the material with the best adsorption performance; inputting the best material into the candidate material library to obtain a synthesis method for the best material.

2. The method for screening adsorbent materials based on machine learning according to claim 1, It is characterized in that In step (2), the machine learning model includes a feedforward neural network, a random forest, a gradient boosting tree, and a multi-granularity cascade forest.

3. The method for screening adsorbent materials based on machine learning according to claim 1, It is characterized in that In step (2), the method for determining the optimal hyperparameters is Gridsearch.

4. The method for screening adsorbent materials based on machine learning according to claim 1, It is characterized in that In step (2), the goodness of fit R 2 The best prediction model was evaluated by R and RMSE. 2 The calculation formula of RMSE is as follows: Where n represents the total number of data sets, and yi are the best model prediction value and true value of the ith data, respectively. All predicted values The average value of .

5. The method for screening adsorbent materials based on machine learning according to claim 1, It is characterized in that In step (2), the Shapley value in feature engineering is used as an indicator to evaluate the contribution of feature variables, and the key parameters affecting the adsorption performance are identified through the retention of the model prediction performance.

6. The method for screening adsorbent materials based on machine learning according to claim 1, It is characterized in that In step (3), if there is no corresponding optimal material synthesis method in the candidate material library, the synthesis method of the optimal material is obtained by obtaining the characteristic information of the metal center and organic ligand of the optimal material.

7. A device for screening adsorbent materials based on machine learning, It is characterized in that Includes a dataset building module to obtain descriptors of materials, pollutants, and adsorption processes, perform data cleaning, and build raw datasets; The model pre-training module is used to generate multiple initial models containing hyperparameters according to different algorithms, and adjust the parameters to select the optimal hyperparameters; wherein, the material descriptor is obtained from the adsorption material information, and the adsorption material information contains a CIF structure file; if the adsorption material information does not contain a CIF structure file, the adsorption material information is processed by Material studio to obtain the CIF structure file; the adsorption material descriptor includes a mass specific surface area descriptor, a volume specific surface area descriptor, a topological structure descriptor, a metal center descriptor and an organic ligand descriptor; the pollutant descriptor includes a molecular molar refractive index descriptor, a molecular dipole / polarizability descriptor, a hydrogen bond proton acceptor ability descriptor, a hydrogen bond proton donor ability descriptor and a molecular volume descriptor; the adsorption process descriptor is the adsorption amount descriptor in the adsorption isotherm; the model construction module is used to select the best prediction model, the most suitable training set test set division ratio and the best sampling method; wherein the optimal sampling method is: point sampling and group sampling evaluation, and the best sampling is selected. The method comprises the following steps: the group sampling is to group the raw data in units of adsorption isotherms, and the point sampling is to group the raw data without grouping; the feature engineering module is used to compress features and identify key parameters that affect model performance and adsorption effect; the test module is used to input a test set to verify the predictive performance of the model; the material screening module is used to screen the best adsorption material for a specific pollutant, wherein a candidate material library is established, and the candidate material library includes the material descriptor; the pollutant to be tested is input into the best prediction model to locate the material with the best adsorption performance; the best material is input into the candidate material library to obtain a synthesis method for the best material; and the guided synthesis module is used to input the best adsorption material and output the synthesis method for the best material.

Citation Information

Patent Citations

  • Building method of modified bioadsorbent structure-activity relationship model and application thereof

    CN104573273A

  • Chemical material adsorption performance prediction method and device based on automatic machine learning

    CN112966447A