MOFs synthesis automation simulation method and system based on large language model

Through the automated simulation method of MOFs synthesis based on large language model, the problems of long experimental cycles and low efficiency in the MOFs synthesis process are solved, and the precise synthesis and performance optimization of MOFs materials are achieved, reducing R&D costs and cycles.

CN120015207APending Publication Date: 2025-05-16UNIV OF SCI & TECH BEIJING

Patent Information

Application Number
CN202510205030.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the MOFs synthesis process, the existing technology has problems such as long experimental cycle, low efficiency, huge resource consumption, and the "black box" of machine learning models, making it difficult to achieve efficient directional synthesis.

Method used

The MOFs synthesis automation simulation method based on large language models is adopted, and the performance of MOFs materials is optimized through natural language command interaction, the data analysis process is simplified, and the performance-optimized MOFs sample synthesis scheme is generated.

Benefits of technology

It realizes the precise synthesis of MOFs materials, reduces R&D costs, shortens R&D cycle, improves analysis efficiency and accuracy, and provides convenient and efficient experimental synthesis tools for scientific researchers without programming code background.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015207A_ABST
    Figure CN120015207A_ABST
Patent Text Reader

Abstract

The invention provides an MOFs synthesis automation simulation method and system based on a large language model, and belongs to the field of artificial intelligence and nanometer material synthesis. The method comprises the following steps: preparing MOFs samples under different synthesis conditions, carrying out structural characterization and performance testing, and generating and storing a synthesis condition-structure-performance data set; a user puts forward a sample synthesis requirement, an LLM generation general task is divided into data field analysis, data correlation analysis, feature engineering processing, machine learning model construction, machine learning model optimization, machine learning model prediction and analysis subtasks, and subtask planning is generated and fed back; after the user confirms, the LLM reads a related MOFs sample data set according to the disassembled subtasks; and according to the subtask planning, calling a code interpreter tool, performing programming and execution of the subtasks, integrating results of all the subtasks, generating a report and outputting an MOFs sample synthesis scheme. According to the invention, the research and development period of MOFs is shortened, and the cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence and nanocomposite catalytic materials, and specifically relates to an automated simulation method and system for MOFs synthesis based on a large language model. Background Art

[0002] Metal-Organic Framework (MOFs) is a crystalline porous material with a periodic network structure assembled by inorganic structural units and organic ligands through coordination bonds. Due to its unique porous structure, adjustable chemical functions and a wide range of topological structures, MOFs have broad application prospects in gas storage, separation, catalysis and sensing. However, the development and optimization of MOFs involve complex principles, requiring researchers with high quality, solid theoretical foundation, rich experimental experience, and the ability to integrate interdisciplinary knowledge (covering organic chemistry, inorganic chemistry, physical chemistry, analytical chemistry, computational chemistry, physics, biology, and materials science and engineering, etc.); in addition, due to the diverse synthesis conditions of MOFs and the complex influences, the exploration of optimal synthesis conditions often relies on a large number of trial and error experiments, which are long, inefficient, and consume a lot of resources. Therefore, the development of a method that can intelligently guide the synthesis process of MOFs has become a key issue that needs to be solved in the field of materials science.

[0003] The rise of machine learning technology has provided a new solution to this challenge. As a powerful data processing and analysis tool, machine learning can reveal the laws and patterns hidden behind the data, capture "chemical intuition" in high-dimensional and complex chemical space, provide accurate predictions and guidance for the synthesis of MOFs, and promote the transformation of MOFs synthesis from experimental trial and error to intelligent synthesis. However, although machine learning methods provide new opportunities for catalytic applications, they still face many challenges in practical applications. On the one hand, high-level programming skills and data processing capabilities are prerequisites for the application of machine learning methods. This requirement constitutes a barrier to application for researchers with non-technical backgrounds to a certain extent; on the other hand, machine learning models are often regarded as "black box" systems, and their internal reasoning processes lack transparency and explainability, making it difficult for researchers to understand and trust the prediction results of the model. More importantly, traditional machine learning mainly relies on numerical methods to predict material properties, and shows obvious limitations when processing unstructured data (such as experimental schemes, literature descriptions, etc.). Such data often contain rich synthetic knowledge and experience, and it is difficult to fully reflect the true impact of synthetic conditions by simply relying on numerical representation. Therefore, exploring artificial intelligence methods that are more convenient, explainable, and can effectively integrate structured and unstructured data is of great significance for promoting the further development of efficient and targeted synthesis of MOFs.

[0004] At present, the Large Language Model (LLM) is a deep learning model trained with a large amount of text data, which is usually used to generate natural language text or understand language text. Although it has made remarkable achievements in natural language processing and represents one of the most advanced artificial intelligence models, as a general artificial intelligence tool, it still lacks expertise in specific scientific fields, especially materials science. This limitation is particularly prominent in the highly specialized field of MOFs synthesis.

[0005] In the existing technology, when LLM is applied to simulate MOFs synthesis, since the training data of LLM often does not cover enough MOFs synthesis information, it faces challenges in predicting the specific impact of synthesis conditions on material structure and performance. The root cause of this problem lies in the mismatch between the versatility of LLM and the professionalism in the field of MOFs synthesis. Summary of the invention

[0006] In order to solve the above problems, the present invention provides an automated module method and system for MOFs synthesis based on a large language model, which optimizes the performance of MOFs materials through natural language command interaction, simplifies the complex process of data analysis, changes the traditional trial-and-error experimental method that relies on prior knowledge, and efficiently guides the precise synthesis of MOFs materials.

[0007] In order to achieve the above object, the technical solution adopted by the embodiment of the present invention is as follows:

[0008] In a first aspect, an embodiment of the present invention provides a method for simulating the synthesis of MOFs based on a large language model, comprising the following steps:

[0009] Step S1, obtaining experimental parameters, structural characterization data and performance test data of MOFs samples prepared under different synthesis conditions, synthesis condition-structure-performance data set, and saving them in a local database;

[0010] Step S2, based on the interactive interface of the large language model LLM, the user proposes the synthesis requirements of the MOFs sample; the synthesis requirements of the MOFs sample include the structural data and / or performance optimization requirements of the MOFs sample;

[0011] Step S3, generating a total task according to the synthesis requirements of the MOFs sample proposed by the user, and performing task decomposition, decomposing the overall goal into subtasks such as data field analysis, data correlation analysis, feature engineering processing, machine learning model construction, machine learning model optimization, machine learning model prediction and analysis, generating subtask plans based on the subtasks, and feeding back the total task, subtasks and subtask plans to the user;

[0012] Step S4, the user receives the overall task, subtasks and subtask planning, and evaluates the task; if the task meets the requirements, the interaction is confirmed to be completed and step S5 is executed; if the task does not meet the requirements, the synthesis requirements of the MOFs sample are updated and the process returns to step S2;

[0013] Step S5, according to the disassembled subtasks, read the synthesis condition-structure-performance data set of the relevant MOFs sample from the local database; according to the subtask planning, call the code interpreter tool to write and execute the subtask program, simulate the synthesis of the MOFs sample, and predict the performance of the simulated synthesized MOFs sample;

[0014] Step S6, integrate the results of all subtasks, generate a standardized data analysis report, and output a performance-optimized MOFs sample synthesis plan.

[0015] As a preferred embodiment of the present invention, step S1 saves the MOFs synthesis condition-structure-performance data set in a local database, specifically including: installing and deploying a MySQL database locally, creating a MOFs_db database, creating a data table structure that meets the corresponding requirements based on the column name and data type of each column of the data set, and uploading the local table data to the data table; connecting to the MySQL database in the local Python environment with the help of the pymysql library, implementing the execution of SQL statements in Python, calling the MOFs synthesis condition-structure-performance data set in the MySQL database, and providing complete transaction support and cursor operations.

[0016] As a preferred embodiment of the present invention, the MOFs sample is Ni@UiO-66(Ce).

[0017] As a preferred embodiment of the present invention, the synthesis conditions in step S1 at least include: reaction temperature, reaction time, amount of sodium borohydride, and amount of nickel nitrate; the structural characterization data of the sample include nickel loading, metal cerium valence, oxygen vacancy content, metallic nickel content, and hydrogen adsorption; the performance test data of the sample include catalytic performance, specifically referring to dicyclopentadiene conversion rate and tetrahydrodicyclopentadiene selectivity.

[0018] As a preferred embodiment of the present invention, the characterization method includes: using an inductively coupled plasma method to calculate the number of nickel atoms in each cerium node Ce6O6(BDC)6 in different Ni@UiO-66(Ce) samples, using a hydrogen programmed temperature adsorption method to calculate the hydrogen adsorption amount of different Ni@UiO-66(Ce) samples, and using an X-ray photoelectron spectroscopy method to calculate the trivalent cerium content through a theoretical formula. , adsorbed oxygen content , metallic nickel content , used to characterize the electronic structure of the material surface as an input feature for machine learning.

[0019] As a preferred embodiment of the present invention, the data field analysis in step S2 includes analyzing the structural data of the MOFs sample, analyzing and decomposing the performance optimization requirements; the data correlation analysis includes analyzing the correlation between the synthesis conditions, structural data and performance of the MOFs sample; the feature engineering processing includes selecting synthesis conditions and generating a synthesis process based on the data correlation analysis, and analyzing the characteristics of each synthesis step; the machine learning model construction includes constructing a machine learning model based on the feature engineering processing results; the machine learning model optimization includes training and optimizing the machine learning model; the machine learning model prediction and analysis includes using the trained machine learning model to output the synthesis simulation process of the MOFs sample.

[0020] As a preferred embodiment of the present invention, step S5 specifically includes: implementing multiple rounds of question-and-answer and review processes for the core objectives of each subtask, LLM in data field parsing, data correlation analysis, feature engineering processing, machine learning model construction, and machine learning model optimization; in the machine learning model prediction and analysis stage, calling the MySQL tool to read data in the local environment, calling the code interpreter tool to automatically write and execute natural language-based Python code, and calling the Google Cloud API tool to summarize and save the question-and-answer results of key questions in each stage.

[0021] As a preferred embodiment of the present invention, in the performance prediction process of step S5, LLM selects one or more of the following machine learning algorithms, including: random forest, kernel ridge regression, support vector machine, K nearest neighbor, neural network, Xgboost, Adaboost and gradient boosting decision tree model, and uses Bayesian optimization to tune hyperparameters.

[0022] As a preferred embodiment of the present invention, step S6 generates a standardized output report, specifically: in the process of multiple rounds of dialogue, LLM summarizes the key questions and answers of each subtask, and writes a complete data analysis report based on these questions and answers.

[0023] In a second aspect, an embodiment of the present invention further provides a MOFs synthesis automation simulation system based on a large language model, the system comprising: a data acquisition module, a local database, an interactive interface based on LLM, a memory storage module, a subtask planning module, an external tool module, an action response module and a result output module; wherein,

[0024] The data acquisition module is used to obtain experimental parameters, structural characterization data and performance test data of MOFs samples prepared under different synthesis conditions, synthesis condition-structure-performance data set, and save them in a local database;

[0025] The local database is used to store the synthesis condition-structure-performance data set of MOFs samples;

[0026] The LLM-based interactive interface is used for users to interact with LLM, and users put forward synthesis requirements of MOFs samples; the synthesis requirements of MOFs samples include structural data and / or performance optimization requirements of MOFs samples; and users are also used to evaluate and confirm the overall task, subtasks and subtask planning;

[0027] The memory storage module is used by LLM to store and manage multi-round dialogue information and key system information to achieve long-term knowledge accumulation and preservation;

[0028] The subtask planning module is used to generate a total task according to the synthesis requirements of the MOFs sample proposed by the user, and to perform task disassembly, disassembling the overall goal into subtasks such as data field analysis, data correlation analysis, feature engineering processing, machine learning model construction, machine learning model optimization, machine learning model prediction and analysis, etc., generating subtask planning based on subtasks, and feeding back the total task, subtasks and subtask planning to the user;

[0029] The external tool module is used to equip LLM with various tools and grant it usage rights to enhance its ability and flexibility in performing tasks;

[0030] The action response module is used to read the synthesis condition-structure-performance data set of the relevant MOFs sample from the local database according to the disassembled subtasks; according to the subtask planning, call the code interpreter tool to write and execute the subtask program, simulate the synthesis of the MOFs sample, and predict the performance of the simulated synthesized MOFs sample;

[0031] The result output module is used to integrate the results of all subtasks, generate a standardized data analysis report, and output a performance-optimized MOFs sample synthesis plan.

[0032] The solution of the embodiment of the present invention has the following beneficial effects:

[0033] The MOFs synthesis automation simulation method and system based on the large language model provided in the embodiment of the present invention, by constructing the MOFs synthesis automation simulation system based on the large language model, reveals the correlation between the synthesis parameters and the catalytic performance of Ni@UiO-66(Ce), accurately identifies the key factors affecting the material performance, provides a high-performance supported Ni@UiO-66(Ce) material precision customization solution, realizes the efficient planning of material synthesis routes in a wide range of chemical space, effectively avoids the waste of resources caused by the frequent "trial and error" process in traditional experiments, thereby reducing R&D costs and accelerating the R&D process; the present invention automates the complex MOFs machine learning data analysis process, decomposes the overall task into a series of operable subtasks through the intelligent disassembly and execution capabilities of the large language model, and automatically calls the corresponding tools and resources to complete the synthesis simulation process. This fully automated analysis process not only simplifies the steps of data analysis, but also improves the efficiency and accuracy of the analysis. At the same time, the present invention provides researchers without a programming code background with an efficient tool for intelligent material synthesis. The system can be driven to complete code writing and execution of fully automated machine learning data analysis through simple natural language instructions. According to the synthetic data analysis and catalytic performance optimization problems raised by the user, the optimized synthesis scheme and experimental steps are automatically generated, providing researchers with a more convenient and efficient experimental synthesis method.

[0034] Of course, it is not necessary to achieve all of the advantages described above at the same time to implement any product or method of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0036] Figure 1 It is a flow chart of the MOFs synthesis automation simulation method based on the large language model in an embodiment of the present invention;

[0037] Figure 2 is a bar chart of the performance evaluation of eight machine learning models for predicting the hydrogenation performance of MOFs in an embodiment of the present invention;

[0038] Figure 3 It is a scatter plot of the performance evaluation of the Xgboost model for predicting the hydrogenation performance of MOFs in the embodiment of the present invention. DETAILED DESCRIPTION

[0039] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. The components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations. It should be noted that the embodiments of the present invention and the features in the embodiments can also be combined with each other without conflict.

[0040] It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. In the description of the present invention, the terms "first", "second", "third", "fourth", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.

[0041] In view of the existing problems in applying LLM to MOFs synthesis simulation, the embodiment of the present invention proposes a MOFs synthesis automation simulation method and system based on a large language model, which aims to achieve full automation of complex data analysis processes in materials science research by integrating advanced artificial intelligence technology, and accurately customize the synthesis scheme of high-performance materials, so as to overcome the shortcomings of long experimental cycle and high experimental cost, high threshold of machine learning code ability, difficulty in learning unstructured data, and no interpretability of the model "black box" when optimizing MOFs material performance. The system includes multiple functional modules such as memory storage, subtask planning, external tools, and action response. Each module works together to optimize MOFs material performance through natural language command interaction, simplify the complex process of data analysis, and obtain rich and in-depth analysis results by providing only the minimum amount of input information. When using, users can interact with the intelligent system to raise MOFs synthesis data analysis and performance optimization problems in natural language. The intelligent system completes the whole process of machine learning data analysis by intention recognition, task decomposition, human-computer interaction modification, and automatic call of corresponding tools to write and execute code. Finally, the intelligent system integrates all subtask results, generates a standardized data analysis report, and outputs a performance-optimized MOFs synthesis scheme. This invention provides researchers with an intelligent synthesis assistant that eliminates code barriers, changes the traditional trial-and-error experimental method that relies on prior knowledge, efficiently guides the precise synthesis of MOFs materials, significantly shortens the MOFs material R&D cycle and reduces costs, and promotes the intelligentization of materials science research.

[0042] like Figure 1 As shown, the MOFs synthesis automation simulation method based on the large language model provided in the embodiment of the present invention comprises the following steps:

[0043] Step S1, obtaining experimental parameters, structural characterization data and performance test data of MOFs samples prepared under different synthesis conditions, synthesis condition-structure-performance data set, and saving them in a local database.

[0044] In this embodiment, the Ni@UiO-66(Ce) catalyzed DCPD hydrogenation reaction is used as an application scenario to illustrate the MOFs synthesis automation simulation method of the present invention. Among them, UiO-66(Ce), as a type of MOFs material, has the characteristics of high coordination number, extremely high specific surface area, uniform microporous structure and outstanding structural stability; nickel, as an excellent non-precious metal catalyst, has strong hydrogenation activity. Loading metal nickel nanoparticles on UiO-66(Ce) can effectively improve the dispersion and utilization rate of active metal components, and can significantly enhance the catalytic activity of the catalyst.

[0045] In this step, Ni@UiO-66(Ce) samples under different synthesis conditions were prepared, the structural characteristics of different samples were evaluated by characterization methods, and the catalytic performance of the samples for dicyclopentadiene (DCPD) hydrogenation reaction was tested. The synthesis conditions, sample structures and sample performance data were collected to form a Ni@UiO-66(Ce) synthesis condition-structure-performance data set, and the Ni@UiO-66(Ce) synthesis condition-structure-performance data set was saved in the local database for all functional modules of LLM to call.

[0046] In this step, the synthesis conditions at least include: reaction temperature, reaction time, amount of sodium borohydride, amount of nickel nitrate, etc. The structural characteristics of the sample include nickel loading, metal cerium valence, oxygen vacancy content, metallic nickel content, hydrogen adsorption, and characterization methods include: using inductively coupled plasma (ICP) method to calculate the number of nickel atoms in each cerium node Ce6O6(BDC)6 in different Ni@UiO-66(Ce) samples, using hydrogen temperature programmed desorption (H2-TPD) method to calculate the hydrogen adsorption of different Ni@UiO-66(Ce) samples, and using X-ray photoelectron spectroscopy (XPS) method to calculate the trivalent cerium content by theoretical formula , adsorbed oxygen content , metallic nickel content , used to characterize the electronic structure of the material surface as an input feature for machine learning. The catalytic performance includes dicyclopentadiene conversion and tetrahydrodicyclopentadiene selectivity. Preferably, during the experiment, the temperature parameter is accurately set to 0.1.

[0047] In this step, the MOFs synthesis condition-structure-performance data set is saved in the local database, specifically including: installing and deploying the MySQL database locally, creating the MOFs_db database, creating a data table structure that meets the corresponding requirements based on the column name and data type of each column of the data set, and uploading the local table data to the data table; connecting to the MySQL database in the local Python environment with the help of the pymysql library, implementing the execution of SQL statements in Python, calling the MOFs synthesis condition-structure-performance data set in the MySQL database, and providing complete transaction support and cursor operations. For example, for the Ni@UiO-66(Ce) synthesis condition-structure-performance data set, the Ni_UiO66_Ce_db database is created to meet the corresponding data table structure.

[0048] Step S2, based on the interactive interface of the large language model LLM, the user proposes the synthesis requirements of the MOFs sample; the synthesis requirements of the MOFs sample include the structural data and / or performance optimization requirements of the MOFs sample.

[0049] In this step, the large language model LLM is any one of an open source large language model and a commercial large language model, and preferably a GPT-4o model such as Open AI. During the experiment, the temperature parameter is precisely set to 0.1.

[0050] Step S3, generates a total task according to the synthesis requirements of the MOFs sample proposed by the user, and decomposes the task, decomposing the overall goal into subtasks such as data field analysis, data correlation analysis, feature engineering processing, machine learning model construction, machine learning model optimization, machine learning model prediction and analysis, etc., generates subtask plans based on the subtasks, and feeds back the total task, subtasks and subtask plans to the user.

[0051] In this step, the data field analysis includes analyzing the structural data of the MOFs sample, analyzing and deconstructing the performance optimization requirements; the data correlation analysis includes analyzing the correlation between the synthesis conditions, structural data and performance of the MOFs sample; the feature engineering processing includes selecting synthesis conditions and generating a synthesis process based on the data correlation analysis, and analyzing the characteristics of each synthesis step; the machine learning model construction includes constructing a machine learning model based on the feature engineering processing results; the machine learning model optimization includes training and optimizing the machine learning model; the machine learning model prediction and analysis includes using the trained machine learning model to output the synthesis simulation process of the MOFs sample.

[0052] In this step, the subtask planning refers to the specific execution steps of each subtask. When executing the subtask planning at each subtask stage, it also includes: using a series of disassembly examples as few-shots to carry out complex intent recognition and complex task decomposition for user needs, and modifying and adjusting the behavior mode and decision-making strategy of the large language model itself through human-computer interaction prompts.

[0053] In step S4, the user receives the overall task, subtasks and subtask planning, and evaluates the task; if the task meets the requirements, the interaction is confirmed to be completed and step S5 is executed; if the task does not meet the requirements, the synthesis requirements of the MOFs sample are updated and the process returns to step S2.

[0054] Step S5, according to the disassembled subtasks, read the synthesis condition-structure-performance data set of the relevant MOFs sample from the local database; according to the subtask planning, call the code interpreter tool to write and execute the subtask program, simulate the synthesis of the MOFs sample, and predict the performance of the simulated synthesized MOFs sample.

[0055] In this step, according to the subtask planning, multiple rounds of question-and-answer and review processes are implemented for the core objectives of each subtask, and subtasks such as data field parsing, data correlation analysis, feature engineering processing, machine learning model construction, and machine learning model optimization are performed in sequence; in each subtask, the MySQL tool is called to read data in the local environment, the code interpreter tool is called to automatically write and execute Python code based on natural language, and the Google Cloud API tool is called to summarize and save the question-and-answer results of key questions of each stage of the subtask.

[0056] In the performance prediction process, LLM selects one or more of the following machine learning algorithms, including: random forest, kernel ridge regression, support vector machine, K nearest neighbor, neural network, eXtreme Gradient Boost (Xgboost), adaptive boost (Adaboost) and gradient boosting decision tree (GBDT), and uses Bayesian optimization to tune hyperparameters. Preferably, for Ni@UiO-66(Ce) samples, Figure 2 and Figure 3 As shown, since what needs to be predicted is the catalytic performance, the optimal models for predicting catalytic performance, such as the Xgboost model and the XGBoost model, are selected for performance prediction. In a specific embodiment of the present invention, when the Xgboost model and the XGBoost model are used for performance prediction, the best prediction performance determination coefficient is obtained on the test set, and the determination coefficient R² reaches 0.83.

[0057] When executing subtask planning in this step, you can also use enhanced mode and developer mode. In enhanced mode, the complex task decomposition process and code debugging function will be automatically enabled to improve the overall performance of LLM. In developer mode, LLM will first confirm with the user whether the text or code is correct, and then choose whether to save or execute it, which can greatly improve the model usability.

[0058] Step S6, integrate the results of all subtasks, generate a standardized data analysis report, and output a performance-optimized MOFs sample synthesis plan.

[0059] In this step, a standardized output report is generated. Specifically, in the process of multiple rounds of dialogue, LLM summarizes the key questions and answers of each subtask and writes a complete data analysis report based on these questions and answers.

[0060] Preferably, in the synthesis simulation of the Ni@UiO-66(Ce) sample, the performance optimization requirements proposed by the user are as follows: saturated hydrogenation of DCPD under mild conditions, that is, 100% DCPD conversion at 80°C; 100% THDCPD selectivity.

[0061] Input the above interactive content into the LLM interface to generate an overall task. The overall task is broken down into subtasks of data field analysis, data correlation analysis, feature engineering processing, machine learning model construction, machine learning model optimization, machine learning model prediction and analysis, and generate corresponding subtask plans.

[0062] After the overall task, subtasks and subtask plans are fed back to the user, all subtask plans are executed after further interaction or direct confirmation by the user; after execution, the results of all subtasks are integrated to generate a standardized data analysis report and output a performance-optimized MOFs sample synthesis plan.

[0063] In the data field parsing subtask stage, perform the following steps:

[0064] Step S101, using the extract_data function to load the Ni_UiO66_Ce_all_data dataset into the current Python environment;

[0065] Step S102, using the python_inter function to execute the Python code to analyze basic statistical information of the data, including sample size, mean, standard deviation, minimum, maximum, 25th, 50th (median) and 75th percentiles, etc.;

[0066] Step S103, drawing a scatter plot to visualize the relationship between various characteristics and catalytic performance, and observing whether there is an intuitive linear relationship or trend between the synthesis conditions, structural characteristics and catalytic performance;

[0067] Step S104, while maintaining the same amount of nickel nitrate solution, analyzing the relationship between the characteristics and performance of the nickel clusters and observing their distribution characteristics.

[0068] By executing the planning steps of the data field parsing subtask, the dataset is fully explored to facilitate subsequent in-depth analysis.

[0069] In the data correlation analysis subtask phase, perform the following steps:

[0070] Step S201, using the extract_data function to import the Ni_UiO66_Ce_all_data dataset into the Python environment.

[0071] Step S202, using the python_inter function to write Python code to calculate the correlation coefficient between each field in the Ni_UiO66_Ce_all_data data set and the catalytic performance.

[0072] Step S203, using the fig_inter function to write Python code for visualization, generating relevant heat maps or scatter plots to illustrate the relationship between various parameters and catalytic performance.

[0073] Step S204, based on the heat map and scatter plot, explain the correlation between the synthesis conditions, structural characteristics and catalytic performance, and analyze the results.

[0074] During the feature engineering subtask phase, the following steps are performed:

[0075] Step S301, use python_inter function to extract Ni_UiO66_Ce_all_data from the database and load it into the local Python environment.

[0076] Step S302: Use the python_inter function to execute Python code to perform data cleaning and preprocessing, including processing missing values, feature selection, and feature scaling.

[0077] Step S303, combining multiple features using methods such as genetic algorithms to create new features that are highly correlated with catalytic performance.

[0078] Step S304, using the python_inter function to save the processed data in the Python environment, to facilitate subsequent modeling and optimization tasks.

[0079] In the machine learning model building subtask stage, perform the following steps:

[0080] Step S401, selection of regression model type: including random forest regression, Adaboost regression, XGboost regression and other models.

[0081] Step S402, model training: using the training set data (X_train_scaled and y_train) to train the selected model.

[0082] Step S403, performance prediction: predicting the test set (X_test_scaled).

[0083] Step S404, model evaluation: Calculate the evaluation indicators of the model, such as mean square error (MSE) and coefficient of determination (R²), and present these results in a visual way.

[0084] Step S405, model selection: select the best performing model and save the model training process and its evaluation results

[0085] Step S406, based on the evaluation result, decide whether to return to step S401 or end the model training.

[0086] In the machine learning model optimization subtask stage, the following steps are performed:

[0087] Step S501, selecting the model with the best performance for further optimization.

[0088] Step S502: Use grid search or Bayesian optimization to adjust model hyperparameters to enhance its performance.

[0089] Step S503, predict the test set and calculate evaluation indicators, such as MSE and r-squared, to evaluate the model performance.

[0090] Step S504, draw a relationship diagram between the actual value and the predicted value to visualize the prediction effect of the model.

[0091] Step S505, analyzing the results of the optimization process and summarizing the conclusions.

[0092] In the machine learning model prediction and analysis subtask stage, perform the following steps:

[0093] Step S601, calculate the feature importance of the model. Using the best training model, XGBoost evaluates the importance of each feature for predictive performance.

[0094] Step S602, analyzing key factors affecting performance. Based on the importance of the features, the key factors affecting catalytic performance are analyzed in depth, and the potential influencing mechanism of catalytic performance is discussed.

[0095] Step S603, establish a wider synthesis space. Based on the available data, a reasonable variable range is determined, and a wider synthesis space for material property prediction is constructed.

[0096] Step S604, applying the model to the new synthetic space. In the newly constructed synthetic space, the best model is used for material property prediction.

[0097] Step S605: Present the results through visualization methods. The prediction results are visualized using appropriate methods to facilitate understanding and further analysis of possible preparation optimization strategies.

[0098] A specific synthesis scheme obtained specifically includes the following contents:

[0099] 1 g of 1,4-benzenedicarboxylic acid (H2BDC) was dissolved in 30 mL of N,N-dimethylformamide (DMF) and the solution was introduced into a glass reactor. 3.65 g of Ce(NO3)6(NH4)2·6H2O was dissolved in 12.5 mL of deionized water and 10 mL of formic acid (HCOOH, 98% concentration) was added. The DMF solution of H2BDC was mixed with the Ce(NO3)6(NH4)2·6H2O aqueous solution in a glass vial. The glass reactor was sealed and the mixture was heated in an oil bath at 100 °C for 15 minutes with stirring; after cooling to room temperature, the pale yellow precipitate in the mother liquor was centrifuged. The precipitate was redispersed in 2 mL of DMF and centrifuged; this washing and centrifugation step was repeated two more times. In order to remove the residual DMF from the product, the solid was washed with 2 mL of acetone and centrifuged. This washing and centrifugation step was repeated two more times. Drying in a vacuum at 80 °C for 12 hours gave a white solid. The synthesis process of Ni@UiO-66(Ce) includes: dissolving 100 mg Ni(NO3)2·6H2O in 100 mL methanol to prepare 1.0 g / L Ni(NO3)2·6H2O methanol solution. Take 10 mL Ni(NO3)2·6H2O methanol solution and 10 mL methanol, add 100 mg of synthesized UiO-66(Ce) to the solution, stir at room temperature for 1 hour, dissolve 16.165 mg NaBH4 in 5 mL methanol, and during the reduction process, quickly add NaBH4 solution dropwise, stir at room temperature for 1 hour, and repeat this reduction step once. Redisperse the precipitate and centrifuge it in methanol three times. The obtained solid is dried in a vacuum at 80 °C for 12 hours to obtain a black solid.

[0100] Based on the above synthesis simulation results, the present invention can further verify the simulation results through experiments. The specific steps are as follows: prepare MOFs samples according to the output synthesis scheme, record the synthesis conditions, and perform structural characterization and performance testing on the synthesized MOFs samples.

[0101] Preferably, taking the Ni@UiO-66(Ce) sample as an example, the Ni@UiO-66(Ce) material is prepared according to the output synthesis scheme, and the prepared Ni@UiO-66(Ce) material is tested for the catalytic performance of dicyclopentadiene hydrogenation. In a specific embodiment, taking the above synthesis scheme as an example, the synthesized Ni@UiO-66(Ce) is tested for the catalytic performance of dicyclopentadiene hydrogenation, specifically: using a one-pot method, in a stainless steel autoclave, with external magnetic stirring and electric heating, 20 mg Ni@UiO-66(Ce) catalyst, 200 mg dicyclopentadiene hydrogenation and 5 mL methanol are loaded into the autoclave and sealed. Blow nitrogen into the entire system at room temperature 2 MPa for 30 min to test the air tightness of the unit. Then, blow N2 and H2 into the sealed autoclave and release, repeat three times, and exhaust the air in the system. Subsequently, the entire system is purged with hydrogen and pressurized to 2 MPa. Stir at 600 rpm and raise the temperature to the reaction temperature (80 °C) for a period of time. After the reaction is completed, collect the product and analyze it using a gas chromatography-mass spectrometry detector. Calculate the DCPD conversion rate based on the carbon balance

[0102] , DHDCPD selectivity and THDCPD selectivity

[0103] .

[0104] Based on the same idea, an embodiment of the present invention also provides an MOFs synthesis automation simulation system based on a large language model, the system comprising: a data acquisition module, a local database, an interactive interface based on LLM, a memory storage module, a subtask planning module, an external tool module, an action response module and a result output module.

[0105] The data acquisition module is used to obtain the experimental parameters, structural characterization data and performance test data of MOFs samples prepared under different synthesis conditions, the synthesis condition-structure-performance data set, and save them in the local database;

[0106] The local database is used to store the synthesis condition-structure-performance data set of MOFs samples;

[0107] The LLM-based interactive interface is used for users to interact with LLM, and users put forward synthesis requirements of MOFs samples; the synthesis requirements of MOFs samples include structural data and / or performance optimization requirements of MOFs samples; and users are also used to evaluate and confirm the overall task, subtasks and subtask planning;

[0108] The memory storage module is used by LLM to store and manage multi-round dialogue information and key system information to achieve long-term knowledge accumulation and preservation;

[0109] The subtask planning module is used to generate a total task according to the synthesis requirements of the MOFs sample proposed by the user, and to perform task disassembly, disassembling the overall goal into subtasks such as data field analysis, data correlation analysis, feature engineering processing, machine learning model construction, machine learning model optimization, machine learning model prediction and analysis, etc., generating subtask planning based on subtasks, and feeding back the total task, subtasks and subtask planning to the user;

[0110] The external tool module is used to equip LLM with various tools and grant it usage rights to enhance its ability and flexibility in performing tasks;

[0111] The action response module is used to read the synthesis condition-structure-performance data set of the relevant MOFs sample from the local database according to the disassembled subtasks; according to the subtask planning, call the code interpreter tool to write and execute the subtask program, simulate the synthesis of the MOFs sample, and predict the performance of the simulated synthesized MOFs sample;

[0112] The result output module is used to integrate the results of all subtasks, generate a standardized data analysis report, and output a performance-optimized MOFs sample synthesis plan.

[0113] As described above and in any possible implementation, the memory storage module is also used to store long-term memory of local expert documents and data dictionaries and interactive analysis functions with controllable memory length to support short-term memory of multiple rounds of dialogue.

[0114] As described above and in any possible implementation, the subtask planning module is also used to perform complex intent recognition and complex task decomposition on user needs by taking a series of disassembly examples as few-shot, and to modify and adjust the behavior patterns and decision-making strategies of the large language model itself through human-computer interaction prompts.

[0115] As described above and in any possible implementation, the external tool module is also used to extract a MySQL relational database for local data, a code interpreter for automatic writing and execution of natural language Python code, and a Google cloud service for real-time aggregation and storage of multiple rounds of question-and-answer results.

[0116] As described above and in any possible implementation, the action response module includes an enhanced mode and a developer mode. In the enhanced mode, the complex task decomposition process and the code debugging function are automatically enabled to improve the overall performance of LLM. In the developer mode, LLM will first confirm with the user whether the text or code is correct, and then choose whether to save or execute it, which can greatly improve the model usability.

[0117] It can be seen from the above technical solutions that the MOFs synthesis automation simulation method and system based on the large language model provided in the embodiment of the present invention, by constructing a fully automated MOFs data analysis intelligent agent system based on the large language model, reveals the correlation between the synthesis parameters and the catalytic performance of Ni@UiO-66(Ce), accurately identifies the key factors affecting the material performance, and provides a high-performance supported Ni@UiO-66(Ce) material precision customization solution, realizes the efficient planning of material synthesis routes in a wide range of chemical space, and effectively avoids the waste of resources caused by the frequent "trial and error" process in traditional experiments, thereby reducing R&D costs and accelerating the R&D process; the present invention automates the complex MOFs machine learning data analysis process, and through the intelligent disassembly and execution capabilities of the large language model, decomposes the overall task into a series of operable subtasks, and automatically calls the corresponding tools and resources to complete them. This fully automated analysis process not only simplifies the steps of data analysis, but also improves the efficiency and accuracy of the analysis. At the same time, the present invention provides researchers without a programming code background with an efficient tool for intelligent material synthesis. The system can be driven to complete code writing and execution of fully automated machine learning data analysis through simple natural language instructions. According to the synthetic data analysis and catalytic performance optimization problems raised by the user, the optimized synthesis scheme and experimental steps are automatically generated, providing researchers with a more convenient and efficient experimental synthesis method.

[0118] The above description is only a preferred embodiment of the present invention and an explanation of the technical principles used. It is not intended to limit the scope of the invention claimed for protection, but only represents the preferred embodiment of the present invention. Those skilled in the art should understand that the scope of the invention involved in the present invention is not limited to the technical solution formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the inventive concept. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present invention.

Claims

1. A method for automating the synthesis of MOFs based on a large language model, characterized in that: The steps include: Step S1, obtaining experimental parameters, structural characterization data and performance test data of metal organic framework compound MOFs samples prepared under different synthesis conditions, synthesis condition-structure-performance data set, and saving them in a local database; Step S2, based on the interactive interface of the large language model LLM, the user proposes the synthesis requirements of the MOFs sample; the synthesis requirements of the MOFs sample include the structural data and / or performance optimization requirements of the MOFs sample; Step S3, generating a total task according to the synthesis requirements of the MOFs sample proposed by the user, and performing task decomposition, decomposing the overall goal into data field analysis, data correlation analysis, feature engineering processing, machine learning model construction, machine learning model optimization, machine learning model prediction and analysis subtasks, generating subtask plans based on the subtasks, and feeding back the total task, subtasks and subtask plans to the user; Step S4, the user receives the overall task, subtasks and subtask planning, and evaluates the task; If the task meets the requirements, the interaction is confirmed to be completed and step S5 is executed; If the task is not met, the synthesis requirements of the MOFs sample are updated and the process returns to step S2; Step S5, according to the disassembled subtasks, read the synthesis condition-structure-performance data set of the relevant MOFs sample from the local database; according to the subtask planning, call the code interpreter tool to write and execute the subtask program, simulate the synthesis of the MOFs sample, and predict the performance of the simulated synthesized MOFs sample; Step S6, integrate the results of all subtasks, generate a standardized data analysis report, and output a performance-optimized MOFs sample synthesis plan.

2. The MOFs synthesis automation simulation method based on a large language model according to claim 1, characterized in that: Step S1 saves the MOFs synthesis condition-structure-performance data set in a local database, specifically including: installing and deploying a MySQL database locally, creating a MOFs_db database, creating a data table structure that meets the corresponding requirements based on the column name and data type of each column of the data set, and uploading the local table data to the data table; connecting to the MySQL database in the local Python environment with the help of the pymysql library, implementing the execution of SQL statements in Python, calling the MOFs synthesis condition-structure-performance data set in the MySQL database, and providing complete transaction support and cursor operations.

3. The MOFs synthesis automation simulation method based on a large language model according to claim 1, characterized in that: The MOFs sample is Ni@UiO-66(Ce).

4. The MOFs synthesis automation simulation method based on a large language model according to claim 3, characterized in that: The synthesis conditions described in step S1 at least include: reaction temperature, reaction time, amount of sodium borohydride, and amount of nickel nitrate; the structural characterization data of the sample include nickel loading, metal cerium valence, oxygen vacancy content, metallic nickel content, and hydrogen adsorption; the performance test data of the sample include catalytic performance, specifically referring to dicyclopentadiene conversion rate and tetrahydrodicyclopentadiene selectivity.

5. The MOFs synthesis automation simulation method based on a large language model according to claim 4, characterized in that: The characterization methods include: using the inductively coupled plasma method to calculate the number of nickel atoms in each cerium node Ce6O6(BDC)6 in different Ni@UiO-66(Ce) samples, using the hydrogen temperature-programmed adsorption method to calculate the hydrogen adsorption amount of different Ni@UiO-66(Ce) samples, and using the X-ray photoelectron spectroscopy method to calculate the trivalent cerium content through a theoretical formula. , adsorbed oxygen content , metallic nickel content , used to characterize the electronic structure of the material surface as an input feature for machine learning.

6. The MOFs synthesis automation simulation method based on a large language model according to claim 1, characterized in that: The data field analysis in step S2 includes analyzing the structural data of the MOFs sample, analyzing and disassembling the performance optimization requirements; the data correlation analysis includes analyzing the correlation between the synthesis conditions, structural data and performance of the MOFs sample; the feature engineering processing includes selecting the synthesis conditions and generating the synthesis process based on the data correlation analysis, and analyzing the characteristics of each synthesis step; The machine learning model construction includes constructing a machine learning model based on the result of feature engineering processing; the machine learning model optimization includes training and optimizing the machine learning model; the machine learning model prediction and analysis includes using the trained machine learning model to output a synthetic simulation process of MOFs samples.

7. The MOFs synthesis automation simulation method based on a large language model according to claim 1, characterized in that: Step S5 specifically includes: implementing multiple rounds of question-and-answer and review processes for the core objectives of each subtask, LLM in data field parsing, data correlation analysis, feature engineering processing, machine learning model construction, and machine learning model optimization; in the machine learning model prediction and analysis stage, calling the MySQL tool to read data in the local environment, calling the code interpreter tool to automatically write and execute natural language-based Python code, and calling the Google Cloud API tool to summarize and save the question-and-answer results of key questions in each stage.

8. The MOFs synthesis automation simulation method based on a large language model according to claim 1, characterized in that: In the performance prediction process of step S5, LLM selects one or more of the following machine learning algorithms, including: random forest, kernel ridge regression, support vector machine, K nearest neighbor, neural network, Xgboost, Adaboost and gradient boosting decision tree model, and uses Bayesian optimization to tune hyperparameters.

9. The MOFs synthesis automation simulation method based on a large language model according to claim 1, characterized in that: Step S6 generates a standardized output report, specifically: in the process of multiple rounds of dialogue, LLM summarizes the key questions and answers of each subtask, and writes a complete data analysis report based on these questions and answers.

10. A MOFs synthesis automation simulation system based on a large language model, characterized in that: The system includes: a data acquisition module, a local database, an interactive interface based on LLM, a memory storage module, a subtask planning module, an external tool module, an action response module and a result output module; wherein, The data acquisition module is used to obtain experimental parameters, structural characterization data and performance test data of metal organic framework compound MOFs samples prepared under different synthesis conditions, synthesis condition-structure-performance data set, and save them in a local database; The local database is used to store the synthesis condition-structure-performance data set of MOFs samples; The LLM-based interactive interface is used for users to interact with LLM, and users put forward synthesis requirements of MOFs samples; the synthesis requirements of MOFs samples include structural data and / or performance optimization requirements of MOFs samples; and users are also used to evaluate and confirm the overall task, subtasks and subtask planning; The memory storage module is used by LLM to store and manage multi-round dialogue information and key system information to achieve long-term knowledge accumulation and preservation; The subtask planning module is used to generate a total task according to the synthesis requirements of the MOFs sample proposed by the user, and to perform task disassembly, disassembling the overall goal into subtasks such as data field analysis, data correlation analysis, feature engineering processing, machine learning model construction, machine learning model optimization, machine learning model prediction and analysis, etc., generating subtask planning based on subtasks, and feeding back the total task, subtasks and subtask planning to the user; The external tool module is used to equip LLM with various tools and grant it usage rights to enhance its ability and flexibility in performing tasks; The action response module is used to read the synthesis condition-structure-performance data set of the relevant MOFs sample from the local database according to the disassembled subtasks; according to the subtask planning, call the code interpreter tool to write and execute the subtask program, simulate the synthesis of the MOFs sample, and predict the performance of the simulated synthesized MOFs sample; The result output module is used to integrate the results of all subtasks, generate a standardized data analysis report, and output a performance-optimized MOFs sample synthesis plan.

Citation Information

Patent Citations

  • Semantic-based deep multi-instance weak supervision text classification method

    CN115563284A

  • Product demand analysis method and device

    CN118550507A

  • Multi-agent cooperation system and method for material science

    CN118737346A

  • Computer-generated content based on text classification, semantic relevance, and activation of deep learning large language models

    US20240062019A1

  • Ai-enhanced simulation and modeling experimentation and control

    US20240348663A1

Cited By

  • Zeolite multi-task synthesis auxiliary method and system based on large language model

    CN120808923A