Automated MOF synthesis simulation method and system based on large language model
Patent Information
- Application Number
- PCT/CN2025/089100
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-24
- Filing Date
- 2025-04-15
- Publication Date
- 2026-08-27
Smart Images

Figure CN2025089100_27082026_PF_FP_ABST
Abstract
Description
An Automated Simulation Method and System for MOF Synthesis Based on Large Language Models
[0001] Cross-references
[0002] This application claims priority to Chinese Patent Application No. 202510205030.4, filed on February 24, 2025, entitled "Automatic Simulation Method and System for MOF Synthesis Based on Large Language Model", the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application belongs to the field of artificial intelligence and nanocomposite catalytic materials, specifically involving an automated simulation method and system for MOF synthesis based on a large language model. Background Technology
[0004] Metal-organic frameworks (MOFs) are crystalline porous materials with periodic network structures assembled from inorganic structural units and organic ligands through coordination bonds. Due to their unique porous structure, tunable chemical functions, and diverse topologies, MOFs hold great promise for applications in gas storage, separation, catalysis, and sensing. However, the development and optimization of MOFs involve complex principles, requiring researchers to be highly qualified, possess a solid theoretical foundation, and have extensive experimental experience. They also need the ability to integrate interdisciplinary knowledge (covering organic chemistry, inorganic chemistry, physical chemistry, analytical chemistry, computational chemistry, physics, biology, and materials science and engineering). Furthermore, the diverse and complex influences of MOF synthesis conditions often necessitate extensive trial-and-error experiments to find optimal synthesis conditions, resulting in long cycles, low efficiency, and enormous resource consumption. Therefore, developing an intelligent method to guide the synthesis process of MOFs has become a critical problem urgently needing to be solved in the field of materials science.
[0005] The rise of machine learning technology has provided a new solution to this challenge. As a powerful data processing and analysis tool, machine learning can reveal the hidden patterns and rules behind data, capture "chemical intuition" in high-dimensional and complex chemical spaces, provide accurate predictions and guidance for the synthesis of MOFs, and promote the transformation of MOF synthesis from experimental trial and error to intelligent synthesis. However, although machine learning methods have provided new opportunities for catalysis applications, they still face many challenges in practical applications. On the one hand, high-level programming skills and data processing capabilities are prerequisites for applying machine learning methods, which to some extent constitutes a barrier to application for researchers without technical backgrounds. On the other hand, machine learning models are often regarded as "black box" systems, and their internal reasoning processes lack transparency and interpretability, making it difficult for researchers to understand and trust the model's prediction results. More importantly, traditional machine learning mainly relies on numerical methods to predict material properties, which shows obvious limitations when dealing with unstructured data (such as experimental schemes, literature descriptions, etc.). This type of data often contains rich synthetic knowledge and experience, and it is difficult to fully reflect the real influence of synthetic conditions by simply relying on numerical representation. Therefore, exploring more convenient, interpretable, and effective artificial intelligence methods that can integrate structured and unstructured data is of great significance for promoting the further development of efficient directed synthesis of MOFs.
[0006] Currently, Large Language Models (LLMs) are deep learning models trained on large amounts of text data, typically used to generate or understand natural language text. While they have achieved significant success in natural language processing and represent one of the most advanced artificial intelligence models available, as a general-purpose AI tool, they still lack expertise in specific scientific fields, particularly materials science. This limitation is especially pronounced in the highly specialized field of MOF synthesis.
[0007] In existing technologies, when using LLM to simulate MOF synthesis, the training data of LLM often does not cover enough MOF synthesis information, which leads to challenges in predicting the specific impact of synthesis conditions on material structure and properties. The root cause of this problem lies in the mismatch between the versatility of LLM and the specialization of the field of MOF synthesis. Summary of the Invention
[0008] Purpose of the invention
[0009] To address the aforementioned issues, this application provides an automated module method and system for MOF synthesis based on a large language model. This method optimizes MOF material properties through natural language command interaction, simplifies the complex data analysis process, changes the traditional trial-and-error experimental method that relies on prior knowledge, and efficiently guides the precise synthesis of MOF materials.
[0010] Solution
[0011] To achieve the above objectives, the technical solution adopted in this application is as follows:
[0012] In a first aspect, embodiments of this application provide an automated simulation method for MOF synthesis based on a large language model, comprising the following steps:
[0013] Step S1: Obtain experimental parameters, structural characterization data, and performance test data of MOFs samples prepared under different synthesis conditions, and save the synthesis condition-structure-performance dataset in a local database.
[0014] Step S2: Based on the interactive interface of the Large Language Model (LLM), the user proposes the synthesis requirements of MOFs samples; the synthesis requirements of MOFs samples include structural data and / or performance optimization requirements of MOFs samples.
[0015] Step S3: Generate the overall task based on the synthesis requirements of MOFs samples proposed by the user, and decompose the task into sub-tasks such as data field parsing, data correlation analysis, feature engineering, machine learning model construction, machine learning model optimization, and machine learning model prediction and analysis. Generate sub-task planning based on the sub-tasks, and feed back the overall task, sub-tasks and sub-task planning to the user.
[0016] In step S4, the user receives the overall task, sub-tasks, and sub-task plans, and evaluates the task; if the task meets the requirements, the user confirms the completion of the interaction and proceeds to step S5; if the task does not meet the requirements, the user updates the synthesis requirements of the MOFs sample and returns to step S2.
[0017] Step S5: Based on the decomposed sub-tasks, read the synthesis conditions-structure-performance dataset of relevant MOFs samples from the local database; according to the sub-task plan, call the code interpreter tool to write and execute the sub-task program, simulate the synthesis of MOFs samples, and predict the performance of the simulated synthesized MOFs samples.
[0018] Step S6: Integrate the results of all subtasks, generate a standardized data analysis report, and output a performance-optimized MOFs sample synthesis scheme.
[0019] As a preferred embodiment of this application, step S1 stores the MOFs synthesis condition-structure-performance dataset in a local database, specifically including: installing and deploying a MySQL database locally, creating the MOFs_db database, creating a data table structure that meets the corresponding requirements based on the column name and data type of each column in the dataset, and uploading the local table data to the data table; using the pymysql library to connect to the MySQL database in the local Python environment, realizing the execution of SQL statements in Python, calling the MOFs synthesis condition-structure-performance dataset in the MySQL database, and providing complete transaction support and cursor operations.
[0020] In a preferred embodiment of this application, the MOF sample is Ni@UiO-66(Ce).
[0021] As a preferred embodiment of this application, the synthesis conditions described in step S1 include at least: reaction temperature, reaction time, sodium borohydride dosage, and nickel nitrate dosage; the structural characterization data of the sample includes nickel loading, cerium valence state, oxygen vacancy content, metallic nickel content, and hydrogen adsorption capacity; the performance test data of the sample includes catalytic performance, specifically dicyclopentadiene conversion and tetrahydrodicyclopentadiene selectivity.
[0022] As a preferred embodiment of this application, the characterization method includes: calculating the number of nickel atoms in each cerium node Ce6O6(BDC)6 of different Ni@UiO-66(Ce) samples using inductively coupled plasma method; calculating the hydrogen adsorption amount of different Ni@UiO-66(Ce) samples using hydrogen temperature-programmed adsorption method; and calculating the trivalent cerium content using X-ray photoelectron spectroscopy method through theoretical formula. Adsorbed oxygen content metallic nickel content Used to characterize the electronic structure of material surfaces, serving as input features for machine learning.
[0023] The above formula W Ce(Ⅲ) In the middle, Ce 3+ Ce 4+ These represent the contents of trivalent and tetravalent cerium ions in the sample, respectively.
[0024] The above formula W Oad In the middle, O ad O represents the adsorbed oxygen content, O=CO represents the lattice oxygen content, and OH represents the adsorbed oxygen content. - This represents the surface oxygen content.
[0025] The above formula W Ni(0) in, Ni 0 Ni represents the content of metallic nickel. 2+This represents the content of divalent nickel ions.
[0026] In a preferred embodiment of this application, the data field parsing in step S2 includes analyzing the structural data of the MOFs sample and analyzing and breaking down the performance optimization requirements; the data correlation analysis includes analyzing the correlation between the synthesis conditions, structural data, and performance of the MOFs sample; the feature engineering process includes selecting synthesis conditions and generating a synthesis process based on the data correlation analysis, and analyzing the characteristics of each synthesis step; the machine learning model construction includes building a machine learning model based on the feature engineering process results; the machine learning model optimization includes training and optimizing the machine learning model; and the machine learning model prediction and analysis includes using the trained machine learning model to output a simulation process for the synthesis of the MOFs sample.
[0027] As a preferred embodiment of this application, step S5 specifically includes: implementing a multi-round question-and-answer and review process for the core objective of each sub-task; in the data field parsing, data correlation analysis, feature engineering, machine learning model construction, machine learning model optimization, and machine learning model prediction and analysis stages, LLM calls the MySQL tool to read data in the local environment, calls the code interpreter tool to automatically write and execute Python code based on natural language, and calls the Google Cloud API tool to summarize and save the question-and-answer results of key questions at each stage.
[0028] In a preferred embodiment of this application, during the performance prediction process in step S5, LLM selects one or more of the following machine learning algorithms, including: random forest, kernel ridge regression, support vector machine, K nearest neighbors, neural network, Xgboost, Adaboost and gradient boosting decision tree model, and uses Bayesian optimization to tune the hyperparameters.
[0029] As a preferred embodiment of this application, step S6 generates a standardized output report, specifically: in the multi-turn dialogue process, LLM summarizes the key questions and answers of each sub-task and compiles a complete data analysis report based on these questions and answers.
[0030] Secondly, embodiments of this application also provide an automated simulation system for MOF synthesis based on a large language model. The system includes: a data acquisition module, a local database, an LLM-based interactive interface, a memory storage module, a subtask planning module, an external tool module, an action response module, and a result output module; wherein,
[0031] The data acquisition module is used to acquire experimental parameters, structural characterization data, and performance test data of MOFs samples prepared under different synthesis conditions, as well as the synthesis condition-structure-performance dataset, and save it in a local database.
[0032] The local database is used to store the synthesis conditions-structure-performance dataset of MOF samples;
[0033] The LLM-based interactive interface is used for users to interact with the LLM, and for users to propose synthesis requirements for MOFs samples. The synthesis requirements for MOFs samples include structural data and / or performance optimization requirements for MOFs samples. It is also used by users to evaluate and confirm the overall task, sub-tasks, and sub-task planning.
[0034] The memory storage module is used for LLM storage and management of multi-turn dialogue information and key system information, so as to realize the accumulation and preservation of long-term knowledge;
[0035] The subtask planning module is used to generate a total task based on the synthesis requirements of MOFs samples proposed by the user, and to decompose the task into subtasks such as data field parsing, data correlation analysis, feature engineering, machine learning model construction, machine learning model optimization, and machine learning model prediction and analysis. Based on the subtasks, a subtask plan is generated, and the total task, subtasks and subtask plans are fed back to the user.
[0036] The external tool module is used to equip the LLM with a variety of tools and grant it usage permissions to enhance its ability and flexibility in performing tasks;
[0037] The action response module is used to read the synthesis conditions-structure-performance dataset of relevant MOFs samples from the local database according to the decomposed sub-tasks; according to the sub-task plan, call the code interpreter tool to write and execute the sub-task program, simulate the synthesis of MOFs samples, and predict the performance of the simulated synthesized MOFs samples.
[0038] The results output module is used to integrate the results of all subtasks, generate a standardized data analysis report, and output a performance-optimized MOFs sample synthesis scheme.
[0039] The solution in this application has the following beneficial effects:
[0040] The automated simulation method and system for MOF synthesis based on a large language model provided in this application reveals the correlation between synthesis parameters and the catalytic performance of Ni@UiO-66(Ce) by constructing an automated simulation system for MOF synthesis based on a large language model. It accurately identifies key factors affecting material performance, provides a precise customization scheme for high-performance supported Ni@UiO-66(Ce) materials, and realizes efficient planning of material synthesis routes in a wide chemical space. It effectively avoids the resource waste caused by the frequent "trial and error" process in traditional experiments, thereby reducing R&D costs and accelerating the R&D process. This application automates the complex MOF machine learning data analysis process. Through the intelligent decomposition and execution capabilities of the large language model, the overall task is decomposed into a series of operable sub-tasks, and the corresponding tools and resources are automatically called to complete the synthesis simulation process. This fully automated analysis process not only simplifies the data analysis steps but also improves the efficiency and accuracy of the analysis. At the same time, this application provides researchers without a programming background with an efficient tool for intelligent material synthesis. The system can be driven to complete fully automated machine learning data analysis by simply using natural language commands. Based on the synthesis data analysis and catalytic performance optimization problems proposed by the user, it automatically generates optimized synthesis schemes and experimental steps, providing researchers with a more convenient and efficient experimental synthesis method.
[0041] Of course, implementing any product or method of this application does not necessarily require achieving all of the advantages described above at the same time. Attached Figure Description
[0042] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0043] Figure 1 is a flowchart of the automated simulation method for MOF synthesis based on a large language model in an embodiment of this application;
[0044] Figure 2 is a bar chart showing the performance evaluation of eight machine learning models for predicting the hydrogenation performance of MOFs in the embodiments of this application;
[0045] Figure 3 is a scatter plot of the performance evaluation of the Xgboost model for predicting the hydrogenation performance of MOFs in the embodiments of this application. Detailed Implementation
[0046] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can also be combined with each other.
[0047] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In the description of this application, the terms "first," "second," "third," "fourth," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0048] To address the problems existing in applying LLM to MOF synthesis simulation, this application proposes an automated simulation method and system for MOF synthesis based on a large language model. The aim is to fully automate the complex data analysis process in materials science research by integrating advanced artificial intelligence technologies, and to precisely customize synthesis schemes for high-performance materials. This overcomes the shortcomings of optimizing MOF material properties, such as long experimental cycles and high costs, high barriers to machine learning coding skills, difficulty in learning from unstructured data, and the lack of interpretability of "black box" models. The system includes multiple functional modules: memory storage, subtask planning, external tools, and action response. These modules work collaboratively, optimizing MOF material properties through natural language command interaction, simplifying the complex data analysis process. Only a minimal amount of input information is required to obtain rich and in-depth analysis results. In use, users can interact with the intelligent agent system to raise MOF synthesis data analysis and performance optimization questions in natural language. The intelligent agent system completes the entire machine learning data analysis process by recognizing intent, decomposing tasks, modifying human-computer interaction, and automatically calling appropriate tools to write and execute code. Ultimately, the intelligent agent system integrates the results of all subtasks, generates a standardized data analysis report, and outputs a performance-optimized MOFs synthesis scheme. This application provides researchers with an intelligent synthesis assistant that eliminates the coding barrier, changes the traditional trial-and-error experimental method that relies on prior knowledge, efficiently guides the precise synthesis of MOFs materials, significantly shortens the MOFs material development cycle and reduces costs, and promotes the intelligentization of materials science research.
[0049] As shown in Figure 1, the automated simulation method for MOF synthesis based on a large language model provided in this application includes the following steps:
[0050] Step S1: Obtain experimental parameters, structural characterization data, and performance test data of MOFs samples prepared under different synthesis conditions, and save the synthesis condition-structure-performance dataset in a local database.
[0051] In this embodiment, the automated simulation method for MOF synthesis based on Ni@UiO-66(Ce) catalyzing the hydrogenation reaction of DCPD is used as an application scenario to illustrate the present application. UiO-66(Ce), as a type of MOF material, possesses characteristics such as high coordination number, extremely high specific surface area, uniform microporous structure, and outstanding structural stability; nickel, as a high-performing non-noble metal catalyst, exhibits strong hydrogenation activity. Loading metallic nickel nanoparticles onto UiO-66(Ce) can effectively improve the dispersion and utilization rate of the active metal component, significantly enhancing the catalytic activity of the catalyst.
[0052] In this step, Ni@UiO-66(Ce) samples under different synthesis conditions are prepared. The structural characteristics of different samples are evaluated by characterization methods, and the catalytic performance of the samples on the hydrogenation reaction of dicyclopentadiene (DCPD) is tested. Data on synthesis conditions, sample structure, and sample performance are collected to form a Ni@UiO-66(Ce) synthesis conditions-structure-performance dataset, which is then saved in a local database for use by all functional modules of LLM.
[0053] In this step, the synthesis conditions include at least: reaction temperature, reaction time, sodium borohydride dosage, and nickel nitrate dosage. The structural characteristics of the sample include nickel loading, cerium valence state, oxygen vacancy content, metallic nickel content, and hydrogen adsorption capacity. Characterization methods include: calculating the number of nickel atoms per cerium node Ce6O6(BDC)6 in different Ni@UiO-66(Ce) samples using inductively coupled plasma (ICP); calculating the hydrogen adsorption capacity of different Ni@UiO-66(Ce) samples using hydrogen temperature programmed desorption (H2-TPD); and calculating the trivalent cerium content using X-ray photoelectron spectroscopy (XPS) through theoretical formulas. Adsorbed oxygen content metallic nickel content Used to characterize the electronic structure of the material surface, serving as input features for machine learning. The catalytic performance includes dicyclopentadiene conversion and tetrahydrodicyclopentadiene selectivity. Preferably, the temperature parameter is precisely set to 0.1 during the experiment.
[0054] The above formula W Ce(Ⅲ) In the middle, Ce 3+ Ce 4+ These represent the contents of trivalent and tetravalent cerium ions in the sample, respectively.
[0055] The above formula In the middle, O ad O represents the adsorbed oxygen content, O=CO represents the lattice oxygen content, and OH represents the adsorbed oxygen content. - This represents the surface oxygen content.
[0056] The above formula W Ni(0) in, Ni 0 Ni represents the content of metallic nickel. 2+ This represents the content of divalent nickel ions.
[0057] In this step, the MOFs synthesis conditional-structure-performance dataset is stored in a local database. Specifically, this includes: installing and deploying a MySQL database locally, creating the `MOFs_db` database, creating a data table structure that meets the corresponding requirements based on the column names and data types of each column in the dataset, and uploading the local table data to the data table; using the `pymysql` library to connect to the MySQL database in the local Python environment, enabling the execution of SQL statements in Python, calling the MOFs synthesis conditional-structure-performance dataset in the MySQL database, and providing full transaction support and cursor operations. For example, for the `Ni@UiO-66(Ce)` synthesis conditional-structure-performance dataset, the `Ni_UiO66_Ce_db` database is created, with a data table structure that meets the corresponding requirements.
[0058] Step S2: Based on the interactive interface of the Large Language Model (LLM), the user proposes the synthesis requirements of MOFs samples; the synthesis requirements of MOFs samples include structural data and / or performance optimization requirements of MOFs samples.
[0059] In this step, the large language model LLM can be any of the open-source large language model or the commercial large language model. It is preferred to use the GPT-4o model from OpenAI. During the experiment, the temperature parameter is precisely set to 0.1.
[0060] Step S3: Generate the overall task based on the user's requirements for synthesizing MOFs samples, and decompose the task into sub-tasks such as data field parsing, data correlation analysis, feature engineering, machine learning model construction, machine learning model optimization, and machine learning model prediction and analysis. Generate a sub-task plan based on the sub-tasks, and feed back the overall task, sub-tasks, and sub-task plans to the user.
[0061] In this step, the data field parsing includes analyzing the structural data of the MOFs samples and analyzing and breaking down the performance optimization requirements; the data correlation analysis includes analyzing the correlation between the synthesis conditions, structural data, and performance of the MOFs samples; the feature engineering process includes selecting synthesis conditions and generating a synthesis process based on the data correlation analysis, and analyzing the characteristics of each synthesis step; the machine learning model construction includes building a machine learning model based on the feature engineering results; the machine learning model optimization includes training and optimizing the machine learning model; and the machine learning model prediction and analysis includes using the trained machine learning model to output a simulation process for the synthesis of MOFs samples.
[0062] In this step, the subtask planning refers to the specific execution steps of each subtask. Executing subtask planning at each subtask stage also includes: using a series of decomposed instances as few-shots to perform complex intent recognition and complex task decomposition of user needs, and modifying and adjusting the behavior patterns and decision-making strategies of the large language model itself through human-computer interaction prompts.
[0063] In step S4, the user receives the overall task, sub-tasks, and sub-task plans, and evaluates the task. If the task meets the requirements, the user confirms the completion of the interaction and proceeds to step S5. If the task does not meet the requirements, the user updates the synthesis requirements for the MOFs sample and returns to step S2.
[0064] Step S5: Based on the decomposed sub-tasks, read the synthesis conditions-structure-performance dataset of relevant MOFs samples from the local database; according to the sub-task plan, call the code interpreter tool to write and execute the sub-task program, simulate the synthesis of MOFs samples, and predict the performance of the simulated synthesized MOFs samples.
[0065] In this step, based on the subtask plan, a multi-round question-and-answer and review process is implemented for the core objectives of each subtask, sequentially executing subtasks such as data field parsing, data correlation analysis, feature engineering, machine learning model building, and machine learning model optimization. In each subtask, MySQL tools are called to read data in the local environment, code interpreter tools are called to automatically write and execute Python code based on natural language, and Google Cloud API tools are called to summarize and save the question-and-answer results of key questions in each stage of the subtask.
[0066] In the performance prediction process, LLM selects one or more of the following machine learning algorithms, including: Random Forest, Kernel Ridge Regression, Support Vector Machine, K-Nearest Neighbors, Neural Network, Extreme Gradient Boost (XGBoost), Adaptive Boost (AdaBoost), and Gradient Boosting Decision Tree (GBDT), and uses Bayesian optimization for hyperparameter tuning. Preferably, for the Ni@UiO-66(Ce) sample, as shown in Figures 2 and 3, since the catalytic performance needs to be predicted, the optimal models for predicting catalytic performance, such as the XGBoost model and XGBoost model, are selected for performance prediction. In a specific embodiment of this application, the XGBoost model and XGBoost model achieved the best predictive performance determination coefficient on the test set, with a determination coefficient R² of 0.83.
[0067] In this step, when planning subtasks, you can also use enhanced mode and developer mode. In enhanced mode, the complex task decomposition process and code debugging function will be automatically enabled to improve the overall performance of LLM. In developer mode, LLM will first confirm with the user whether the text or code is correct before choosing whether to save or execute, which can greatly improve the usability of the model.
[0068] Step S6: Integrate the results of all subtasks, generate a standardized data analysis report, and output a performance-optimized MOFs sample synthesis scheme.
[0069] In this step, a standardized output report is generated. Specifically, in the multi-turn dialogue process, LLM summarizes the key questions and answers of each subtask and compiles a complete data analysis report based on these questions and answers.
[0070] Preferably, in the synthesis simulation of Ni@UiO-66(Ce) sample, the performance optimization requirements proposed by the user are as follows: DCPD saturated hydrogenation under mild conditions, i.e., DCPD conversion of 100% at 80℃; THDCPD selectivity of 100%.
[0071] Input the above interactive content into the LLM interface to generate the overall task. The overall task is broken down into subtasks such as data field parsing, data correlation analysis, feature engineering, machine learning model building, machine learning model optimization, and machine learning model prediction and analysis, and corresponding subtask plans are generated.
[0072] After the overall task, sub-tasks, and sub-task plans are fed back to the user, and the user further interacts or directly confirms, all sub-task plans are executed. After execution, the results of all sub-tasks are integrated to generate a standardized data analysis report and output a performance-optimized MOFs sample synthesis scheme.
[0073] During the data field parsing subtask phase, perform the following steps:
[0074] Step S101: Use the extract_data function to load the Ni_UiO66_Ce_all_data dataset into the current Python environment;
[0075] Step S102: Use the python_inter function to execute Python code and analyze the basic statistical information of the data, including sample size, mean, standard deviation, minimum, maximum, 25th, 50th (median), and 75th percentiles, etc.
[0076] Step S103: Draw a scatter plot to visualize the relationship between various features and catalytic performance, and observe whether there is an intuitive linear relationship or trend between synthesis conditions, structural features and catalytic performance.
[0077] Step S104: While maintaining a consistent amount of nickel nitrate solution, analyze the relationship between the characteristics and properties of the nickel clusters and observe their distribution characteristics.
[0078] By executing the planned steps of the data field parsing subtask, the dataset is fully explored to facilitate subsequent in-depth analysis.
[0079] In the data correlation analysis subtask phase, perform the following steps:
[0080] Step S201: Use the extract_data function to import the Ni_UiO66_Ce_all_data dataset into the Python environment.
[0081] Step S202: Use the python_inter function to write Python code to calculate the correlation coefficient between each field in the Ni_UiO66_Ce_all_data dataset and the catalytic performance.
[0082] Step S203: Use the fig_inter function to write Python code for visualization, generating relevant heatmaps or scatter plots to illustrate the relationship between various parameters and catalytic performance.
[0083] Step S204: Based on the heatmap and scatter plot, interpret the correlation between synthesis conditions, structural characteristics and catalytic performance, and analyze the results.
[0084] During the feature engineering subtask phase, the following steps are performed:
[0085] Step S301: Use the python_inter function to extract Ni_UiO66_Ce_all_data from the database and load it into the local Python environment.
[0086] Step S302 involves using the `python_inter` function to execute Python code for data cleaning and preprocessing. This includes operations such as handling missing values, feature selection, and feature scaling.
[0087] Step S303: Use methods such as genetic algorithms to combine multiple features to create new features that are highly correlated with catalytic performance.
[0088] Step S304: Use the python_inter function to save the processed data in the Python environment for subsequent modeling and optimization tasks.
[0089] During the machine learning model building subtask phase, the following steps are performed:
[0090] Step S401, selection of regression model type: including models such as random forest regression, Adaboost regression, and XGboost regression.
[0091] Step S402, Model Training: Train the selected model using the training set data (X_train_scaled and y_train).
[0092] Step S403, Performance prediction: Make predictions for the test set (X_test_scaled).
[0093] Step S404, Model Evaluation: Calculate the evaluation metrics of the model, such as mean squared error (MSE) and coefficient of determination (R2), and present these results in a visual manner.
[0094] Step S405, Model Selection: Select the best-performing model and save the model training process and its evaluation results.
[0095] Step S406: Based on the evaluation results, decide whether to return to step S401 or end model training.
[0096] During the machine learning model optimization subtask phase, the following steps are performed:
[0097] Step S501: Select the model with the best performance for further optimization.
[0098] Step S502: Use methods such as grid search or Bayesian optimization to adjust the model hyperparameters to enhance its performance.
[0099] Step S503: Make predictions on the test set and calculate evaluation metrics, such as MSE and r-squared, to evaluate model performance.
[0100] Step S504: Draw a graph showing the relationship between the actual value and the predicted value to visualize the model's prediction effect.
[0101] Step S505: Analyze the results of the optimization process and summarize the conclusions.
[0102] In the machine learning model prediction and analysis subtask phase, the following steps are performed:
[0103] Step S601: Calculate the feature importance of the model. Use the best training mode. XGBoost evaluates the importance of each feature for predictive performance.
[0104] Step S602: Analyze the key factors affecting performance. Based on the importance of the characteristics, the key factors affecting catalytic performance were analyzed in depth, and the potential mechanisms affecting catalytic performance were discussed.
[0105] Step S603: Establish a broader synthesis space. Based on the available data, a reasonable range of variables was determined, and a broader synthesis space for predicting material properties was constructed.
[0106] Step S604: Apply the model to the new synthesis space. In the newly constructed synthesis space, the best model is used for material property prediction.
[0107] Step S605: Present the results using visualization methods. Appropriate visualization methods are used to facilitate understanding and further analysis of possible optimization strategies.
[0108] The specific synthetic scheme obtained includes the following:
[0109] 1 g of 1,4-phenylenediic acid (H₂BDC) was dissolved in 30 mL of N,N-dimethylformamide (DMF), and the solution was introduced into a glass reactor. 3.65 g of Ce(NO₃)₆(NH₄)₂·6H₂O was dissolved in 12.5 mL of deionized water, and 10 mL of formic acid (HCOOH, 98% concentration) was added. The H₂BDC DMF solution was mixed with the Ce(NO₃)₆(NH₄)₂·6H₂O aqueous solution in a glass vial. The glass reactor was sealed, and the mixture was heated and stirred in an oil bath at 100 °C for 15 minutes. After cooling to room temperature, the pale yellow precipitate in the mother liquor was centrifuged. The precipitate was redispersed in 2 mL of DMF and centrifuged; this washing and centrifugation step was repeated at least twice. To remove residual DMF from the product, the solid was washed with 2 mL of acetone and centrifuged. This washing and centrifugation step was repeated twice. The solid was dried under vacuum at 80 °C for 12 hours to obtain a white solid. The synthesis process of Ni@UiO-66(Ce) includes: dissolving 100 mg of Ni(NO3)2·6H2O in 100 mL of methanol to prepare a 1.0 g / L Ni(NO3)2·6H2O methanol solution. Take 10 mL of the Ni(NO3)2·6H2O methanol solution and 10 mL of methanol, add 100 mg of the synthesized UiO-66(Ce) to the solution, stir at room temperature for 1 hour, then dissolve 16.165 mg of NaBH4 in 5 mL of methanol. During the reduction process, rapidly add the NaBH4 solution dropwise, stir at room temperature for 1 hour, and repeat this reduction step once. Redisperse the precipitate and centrifuge three times in methanol. The resulting solid is dried under vacuum at 80 °C for 12 hours to obtain a black solid.
[0110] Based on the above synthesis simulation results, this application can further verify the simulation results through experiments. The specific steps are as follows: prepare MOF samples according to the output synthesis scheme, record the synthesis conditions, and perform structural characterization and performance testing on the synthesized MOF samples.
[0111] Preferably, taking the Ni@UiO-66(Ce) sample as an example, Ni@UiO-66(Ce) material is prepared according to the output synthesis scheme, and the catalytic performance of the prepared Ni@UiO-66(Ce) material for dicyclopentadiene hydrogenation is tested. In a specific embodiment, taking the above synthesis scheme as an example, the catalytic performance of the synthesized Ni@UiO-66(Ce) for dicyclopentadiene hydrogenation is tested as follows: using a one-pot method, in a stainless steel autoclave with external magnetic stirring and electric heating, 20 mg of Ni@UiO-66(Ce) catalyst, 200 mg of dicyclopentadiene hydrogenation, and 5 mL of methanol are loaded into the autoclave and sealed. Nitrogen gas is purged into the entire system at room temperature and 2 MPa for 30 min, and the air tightness of the unit is tested. Then, N2 and H2 are purged into the sealed autoclave and released, repeated three times to remove air from the system. Subsequently, the entire system is purged with hydrogen gas and pressurized to 2 MPa. The mixture was stirred at 600 rpm, and the temperature was increased to the reaction temperature (80°C) for a period of time. After the reaction was complete, the product was collected and analyzed using gas chromatography-mass spectrometry (GC-MS). The DCPD conversion rate was calculated based on the carbon balance. DHDCPD Selectivity and THDCPD selectivity
[0112] Based on the same idea, this application also provides an automated simulation system for MOF synthesis based on a large language model. The system includes: a data acquisition module, a local database, an LLM-based interactive interface, a memory storage module, a subtask planning module, an external tool module, an action response module, and a result output module.
[0113] The data acquisition module is used to acquire experimental parameters, structural characterization data and performance test data of MOFs samples prepared under different synthesis conditions, and to store the synthesis condition-structure-performance dataset in a local database.
[0114] The local database is used to store the synthesis conditions-structure-performance dataset of MOF samples;
[0115] The LLM-based interactive interface is used for users to interact with the LLM, and for users to propose synthesis requirements for MOFs samples. The synthesis requirements for MOFs samples include structural data and / or performance optimization requirements for MOFs samples. It is also used by users to evaluate and confirm the overall task, sub-tasks, and sub-task planning.
[0116] The memory storage module is used for LLM storage and management of multi-turn dialogue information and key system information, so as to realize the accumulation and preservation of long-term knowledge;
[0117] The subtask planning module is used to generate a total task based on the synthesis requirements of MOFs samples proposed by the user, and to decompose the task into subtasks such as data field parsing, data correlation analysis, feature engineering, machine learning model construction, machine learning model optimization, and machine learning model prediction and analysis. Based on the subtasks, a subtask plan is generated, and the total task, subtasks and subtask plans are fed back to the user.
[0118] The external tool module is used to equip the LLM with a variety of tools and grant it usage permissions to enhance its ability and flexibility in performing tasks;
[0119] The action response module is used to read the synthesis conditions-structure-performance dataset of relevant MOFs samples from the local database according to the decomposed sub-tasks; according to the sub-task plan, call the code interpreter tool to write and execute the sub-task program, simulate the synthesis of MOFs samples, and predict the performance of the simulated synthesized MOFs samples.
[0120] The results output module is used to integrate the results of all subtasks, generate a standardized data analysis report, and output a performance-optimized MOFs sample synthesis scheme.
[0121] As described above and in any possible implementation, the memory storage module is also used to store long-term memory of local expert documents and data dictionaries, and short-term memory of interactive analysis functions with controllable memory length to support multi-turn dialogues.
[0122] As described above and in any possible implementation, the subtask planning module is also used to perform complex intent recognition and complex task decomposition of user needs by using a series of decomposed instances as few-shots, and to modify and adjust the behavior patterns and decision-making strategies of the large language model itself through human-computer interaction prompts.
[0123] As described above and in any possible implementation, the external tool module is also used to extract local data from a MySQL relational database, a code interpreter for automatically writing and executing Python code based on natural language, and a Google Cloud service for real-time aggregation and storage of multi-round question-and-answer results.
[0124] As described above and in any possible implementation, the action response module includes an enhanced mode and a developer mode. In the enhanced mode, the complex task decomposition process and code debugging function are automatically enabled to improve the overall performance of the LLM. In the developer mode, the LLM will first confirm with the user whether the text or code is correct before choosing whether to save or execute it, which can greatly improve the usability of the model.
[0125] As can be seen from the above technical solutions, the automated simulation method and system for MOF synthesis based on a large language model provided in this application reveals the correlation between synthesis parameters and the catalytic performance of Ni@UiO-66(Ce) by constructing a fully automated MOF data analysis intelligent system based on a large language model. It accurately identifies the key factors affecting material performance, provides a precise customization scheme for high-performance supported Ni@UiO-66(Ce) materials, realizes efficient planning of material synthesis routes in a wide chemical space, effectively avoids the resource waste caused by the frequent "trial and error" process in traditional experiments, thereby reducing R&D costs and accelerating the R&D process. This application automates the complex MOF machine learning data analysis process. Through the intelligent decomposition and execution capabilities of the large language model, the overall task is decomposed into a series of operable sub-tasks, and the corresponding tools and resources are automatically called to complete them. This fully automated analysis process not only simplifies the data analysis steps but also improves the efficiency and accuracy of the analysis. At the same time, this application provides researchers without a programming background with an efficient tool for intelligent material synthesis. The system can be driven to complete fully automated machine learning data analysis by simply using natural language commands. Based on the synthesis data analysis and catalytic performance optimization problems proposed by the user, it automatically generates optimized synthesis schemes and experimental steps, providing researchers with a more convenient and efficient experimental synthesis method.
[0126] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed, and is not intended to limit the scope of the claimed application, but merely to illustrate preferred embodiments of this application. Those skilled in the art should understand that the scope of the invention involved in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the inventive concept. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application. Industrial applicability
[0127] This application provides an automated simulation method and system for MOF synthesis based on a large language model. The method includes: preparing MOF samples under different synthesis conditions, performing structural characterization and performance testing, generating and saving a synthesis condition-structure-performance dataset; the user submits sample synthesis requirements, and the LLM generates a total task and breaks it down into subtasks: data field parsing, data correlation analysis, feature engineering, machine learning model construction, machine learning model optimization, and machine learning model prediction and analysis; generating and providing feedback on the subtask plan; after user confirmation, the LLM reads the relevant MOF sample dataset according to the decomposed subtasks; based on the subtask plan, it calls a code interpreter tool to write and execute the subtask programs, integrates the results of all subtasks, generates a report, and outputs the MOF sample synthesis scheme.
Claims
1. An automated simulation method for MOF synthesis based on a large language model, characterized in that, Includes the following steps: Step S1: Obtain experimental parameters, structural characterization data, and performance test data of metal-organic framework (MOF) samples prepared under different synthesis conditions, and save the synthesis condition-structure-performance dataset in a local database. Step S2: Based on the interactive interface of the Large Language Model (LLM), the user proposes the synthesis requirements of MOFs samples; the synthesis requirements of MOFs samples include structural data and / or performance optimization requirements of MOFs samples. Step S3: Generate the overall task based on the user's requirements for synthesizing MOFs samples, and decompose the task into sub-tasks: data field parsing, data correlation analysis, feature engineering, machine learning model construction, machine learning model optimization, and machine learning model prediction and analysis. Generate a sub-task plan based on the sub-tasks, and feed back the overall task, sub-tasks, and sub-task plans to the user. In step S4, the user receives the overall task, subtasks, and subtask plans, and evaluates the task. If the task is completed, confirm the interaction is finished and proceed to step S5; If the task is not completed, update the synthesis requirements for the MOFs sample and return to step S2; Step S5: Based on the decomposed sub-tasks, read the synthesis conditions-structure-performance dataset of relevant MOFs samples from the local database; according to the sub-task plan, call the code interpreter tool to write and execute the sub-task program, simulate the synthesis of MOFs samples, and predict the performance of the simulated synthesized MOFs samples. Step S6: Integrate the results of all subtasks, generate a standardized data analysis report, and output a performance-optimized MOFs sample synthesis scheme.
2. The automated simulation method for MOF synthesis based on a large language model according to claim 1, characterized in that, Step S1 saves the MOFs synthesis condition-structure-performance dataset to a local database. Specifically, this includes: installing and deploying a MySQL database locally, creating the MOFs_db database, creating a data table structure that meets the corresponding requirements based on the column name and data type of each column in the dataset, and uploading the local table data to the data table; using the pymysql library to connect to the MySQL database in the local Python environment, enabling the execution of SQL statements in Python, calling the MOFs synthesis condition-structure-performance dataset in the MySQL database, and providing full transaction support and cursor operations.
3. The automated simulation method for MOF synthesis based on a large language model according to claim 1, characterized in that, The MOF sample was Ni@UiO-66(Ce).
4. The automated simulation method for MOF synthesis based on a large language model according to claim 3, characterized in that, The synthesis conditions described in step S1 include at least: reaction temperature, reaction time, sodium borohydride dosage, and nickel nitrate dosage; the structural characterization data of the sample includes nickel loading, cerium valence state, oxygen vacancy content, metallic nickel content, and hydrogen adsorption capacity; the performance test data of the sample includes catalytic performance, specifically the dicyclopentadiene conversion rate and tetrahydrodicyclopentadiene selectivity.
5. The automated simulation method for MOF synthesis based on a large language model according to claim 4, characterized in that, Characterization methods included: calculating the number of nickel atoms in each cerium node Ce6O6(BDC)6 of different Ni@UiO-66(Ce) samples using inductively coupled plasma method; calculating the hydrogen adsorption amount of different Ni@UiO-66(Ce) samples using temperature-programmed hydrogen adsorption method; and calculating the trivalent cerium content using X-ray photoelectron spectroscopy with theoretical formula. Adsorbed oxygen content metallic nickel content Used to characterize the electronic structure of material surfaces, serving as input features for machine learning.
6. The automated simulation method for MOF synthesis based on a large language model according to claim 1, characterized in that, The data field parsing in step S2 includes analyzing the structural data of MOFs samples and analyzing and breaking down the performance optimization requirements; the data correlation analysis includes analyzing the correlation between the synthesis conditions, structural data and performance of MOFs samples; the feature engineering process includes selecting synthesis conditions and generating a synthesis process based on the data correlation analysis, and analyzing the characteristics of each synthesis step. The machine learning model construction includes building a machine learning model based on the results of feature engineering; the machine learning model optimization includes training and optimizing the machine learning model; the machine learning model prediction and analysis includes using the synthesized simulation process of MOFs samples output by the trained machine learning model.
7. The automated simulation method for MOF synthesis based on a large language model according to claim 1, characterized in that, Step S5 specifically includes: implementing a multi-round question-and-answer and review process for the core objectives of each subtask; in the data field parsing, data correlation analysis, feature engineering, machine learning model building, machine learning model optimization, and machine learning model prediction and analysis stages, LLM calls MySQL tools to read data in the local environment, calls code interpreter tools to automatically write and execute Python code based on natural language, and calls Google Cloud API tools to summarize and save the question-and-answer results of key questions at each stage.
8. The automated simulation method for MOF synthesis based on a large language model according to claim 1, characterized in that, In the performance prediction process of step S5, LLM selects one or more of the following machine learning algorithms, including: random forest, kernel ridge regression, support vector machine, K nearest neighbors, neural network, Xgboost, Adaboost and gradient boosting decision tree model, and uses Bayesian optimization to tune hyperparameters.
9. The automated simulation method for MOF synthesis based on a large language model according to claim 1, characterized in that, Step S6 generates a standardized output report, specifically: In the multi-turn dialogue process, LLM summarizes the key questions and answers of each subtask and compiles a complete data analysis report based on these questions and answers.
10. An automated simulation system for MOF synthesis based on a large language model, characterized in that, The system includes: a data acquisition module, a local database, an LLM-based interactive interface, a memory storage module, a subtask planning module, an external tool module, an action response module, and a result output module; among which, The data acquisition module is used to acquire experimental parameters, structural characterization data, and performance test data of metal-organic framework (MOF) samples prepared under different synthesis conditions, as well as a synthesis condition-structure-performance dataset, and save it in a local database. The local database is used to store the synthesis conditions-structure-performance dataset of MOF samples; The LLM-based interactive interface is used for users to interact with the LLM, and for users to propose synthesis requirements for MOFs samples. The synthesis requirements for MOFs samples include structural data and / or performance optimization requirements for MOFs samples. It is also used by users to evaluate and confirm the overall task, sub-tasks, and sub-task planning. The memory storage module is used for LLM storage and management of multi-turn dialogue information and key system information, so as to realize the accumulation and preservation of long-term knowledge; The subtask planning module is used to generate a total task based on the synthesis requirements of MOFs samples proposed by the user, and to decompose the task into subtasks such as data field parsing, data correlation analysis, feature engineering, machine learning model construction, machine learning model optimization, and machine learning model prediction and analysis. Based on the subtasks, a subtask plan is generated, and the total task, subtasks and subtask plans are fed back to the user. The external tool module is used to equip the LLM with a variety of tools and grant it usage permissions to enhance its ability and flexibility in performing tasks; The action response module is used to read the synthesis conditions-structure-performance dataset of relevant MOFs samples from the local database according to the decomposed sub-tasks; according to the sub-task plan, call the code interpreter tool to write and execute the sub-task program, simulate the synthesis of MOFs samples, and predict the performance of the simulated synthesized MOFs samples. The results output module is used to integrate the results of all subtasks, generate a standardized data analysis report, and output a performance-optimized MOFs sample synthesis scheme.