A method for constructing a molecular sieve vertical field intelligent system based on a large language model
By using deep domain training and reinforcement learning, a vertical domain intelligent system for molecular sieves based on a large language model is constructed. This solves the problems of lack of professional knowledge and computational time in general models, and achieves rapid response and high-precision prediction and design of molecular sieve properties, which has broad prospects for material research and development applications.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- EAST CHINA NORMAL UNIV
- Filing Date
- 2026-03-17
- Publication Date
- 2026-06-05
AI Technical Summary
In existing technologies, general-purpose large language models lack expertise in the field of molecular sieves, resulting in high system complexity, long computation time, low training efficiency, and an inability to achieve rapid response and high-precision prediction and design of molecular sieve properties.
We adopt a single strong model architecture, internalize molecular sieve expertise into a large language model through deep domain training, use the GRPO algorithm combined with a multi-dimensional reward function for reinforcement learning, construct a GNN graph neural network agent model, design a "fast first, slow later" tool calling strategy, integrate fast prediction and high-precision calculation tools, and achieve asynchronous calculation scheduling and multi-level caching optimization.
It enables knowledge-based question answering, property prediction, reverse design, and synthesis planning in the field of molecular sieves, improves training efficiency, achieves millisecond-level property prediction, balances response speed and computational accuracy, and reduces system maintenance costs.
Smart Images

Figure CN122154878A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of molecular sieve technology, specifically a method for constructing a vertical domain intelligent system for molecular sieves based on a large language model. Background Technology
[0002] Molecular sieves are a class of crystalline materials with uniform microporous structures, widely used in catalysis, adsorption separation, and ion exchange. Traditional molecular sieve development relies heavily on extensive experimental trial and error and experience accumulation, resulting in long development cycles and high costs. With the development of artificial intelligence technology, deeply integrating large language models with domain knowledge to build intelligent assistants for vertical fields has become an important way to accelerate materials research and development.
[0003] Existing AI-assisted material development methods mainly suffer from the following problems: 1) The general language model lacks professional knowledge in the field of molecular sieves and cannot accurately answer professional questions involving molecular sieve structure, properties, synthesis, etc. It has obvious deficiencies in understanding professional terminology and grasping the conceptual system.
[0004] 2) Existing methods have high system complexity, and the coordination and information transmission between modules are prone to error accumulation, and the maintenance cost is high.
[0005] 3) Prediction of molecular sieve properties usually relies on high-precision first-principles calculations such as density functional theory (DFT) or molecular simulation methods such as grand canonical Monte Carlo (GCMC) and molecular dynamics (MD). These calculation methods are time-consuming and cannot meet the needs of real-time interaction, thus limiting the response speed and user experience of intelligent systems.
[0006] 4) Existing methods typically require the use of real scientific computing software to obtain reward signals during the reinforcement learning training phase, resulting in extremely low training efficiency and making it difficult to achieve large-scale policy optimization.
[0007] 5) The lack of effective tool invocation strategies makes it impossible to achieve a good balance between response speed and calculation accuracy. Either the response is slow, affecting the user experience, or the accuracy is low, affecting the reliability of the results. Summary of the Invention
[0008] The purpose of this invention is to address the shortcomings of existing technologies by proposing a method for constructing a vertical intelligent system for molecular sieves based on a large language model. This method employs a single strong model architecture and internalizes molecular sieve expertise through deep domain training. It injects domain knowledge and tool call specifications through SFT-supervised fine-tuning, and uses the GRPO algorithm combined with a multi-dimensional reward function for reinforcement learning training to enhance reasoning and design capabilities. A GNN (Graph Neural Network) surrogate model is constructed to achieve millisecond-level property prediction, serving as a reward signal source for reinforcement learning and significantly improving training efficiency. Based on the native Function Call capability of the large language model, a tool system is designed, employing a "fast-then-slow" tool call strategy. The GNN surrogate model is prioritized for rapid screening, and high-precision calculations are only triggered after value is confirmed, effectively balancing response speed and computational accuracy. This invention can realize four core functions in the field of molecular sieves: knowledge question answering, property prediction, reverse design, and synthesis planning. The method is highly innovative, the system is complete, and it has broad application prospects in materials research and development.
[0009] The objective of this invention is achieved as follows: a method for constructing an intelligent system for the vertical field of molecular sieves based on a large language model. This method is characterized by first constructing a multi-source heterogeneous data processing flow for the molecular sieve field, acquiring data from multiple sources such as structural databases, literature data, and computational data, achieving structural data standardization and literature information extraction; then, through supervised fine-tuning, injecting domain knowledge and tool call specifications into the large language model, enabling the model to master the professional terminology and conceptual system of the molecular sieve field and learn to output Function Calls according to standardized formats; next, using the GRPO algorithm combined with format compliance rewards, chemical logic rewards, and performance prediction rewards for reinforcement learning training; simultaneously, constructing a GNN graph neural network surrogate model to achieve millisecond-level prediction and uncertainty quantification of molecular sieve properties; designing a tool system based on the native Function Call capabilities of the large language model, integrating fast prediction tools, literature retrieval tools, structure generation tools, and high-precision calculation tools; and finally, constructing a system integration architecture to achieve asynchronous computation scheduling and multi-level caching optimization. The method specifically includes the following steps: 1) Construct a multi-source heterogeneous data processing workflow in the field of molecular sieves to achieve structural data standardization and literature information extraction. 1.1: Obtain official structure data for 250+ topological types from the International Zeolite Association (IZA) structure database, obtain 14,000+ computationally ready MOF structure data from the computationally ready metal-organic framework (CoRE) MOF database, obtain hypothetical structure data from the hypothetical crystal structure (PCOD) database, and obtain organic / organic metal crystal data from the Cambridge Crystal Structures (CSD) database. 1.2: Obtain molecular sieve-related papers from publishers such as Elsevier, Springer, and ACS; obtain preprints in the field of materials science from the academic preprint platform arXiv; and obtain molecular sieve synthesis patents from the USPTO, EPO, and CNIPA patent databases of the United States Patent and Trademark Office, the European Patent Office, and the China National Intellectual Property Administration. 1.3: Obtain material property data from DFT calculations from the Materials Genome Computational Database Materials Project, and computational materials science data from the Computational Materials Science Database NOMAD, to construct a giant canonical Monte Carlo (GCMC) and molecular dynamics (MD) simulation dataset; 1.4: All crystal structures are uniformly converted into three standardized formats: CIF format for storing complete crystallographic information, SMILES-like topological encoding for easy model processing, and node-edge adjacency matrix graph representation for GNN input; 1.5: Natural language processing techniques were used to extract structured information from the literature, including the type of organic structure-directing agent OSDA in the synthesis formulation, the silicon-to-aluminum ratio, water volume, temperature and time, the specific surface area, pore volume, adsorption capacity and acid content in the property data, and experimental phenomena such as the description of the crystallization process and the characterization results.
[0010] 2) Construct a GNN (Graph Neural Network) surrogate model to achieve millisecond-level prediction and uncertainty quantification of molecular sieve properties. 2.1: The MEGNet graph neural network architecture is adopted as the basic architecture, and customized improvements are made for the characteristics of molecular sieves; 2.2: Existing computational data were extracted from databases such as the Materials Project and NOMAD, and experimental measurements were extracted from papers. The molecular simulation software RASPA and density functional theory DFT calculations were used to supplement the training set for key structures, and structure-property pair training data were constructed. 2.3: Train the GNN model to predict the following four types of properties: stability indices represented by formation energy and decomposition temperature, adsorption properties represented by Henry's constant, heat of adsorption and selectivity coefficient, mass transfer properties represented by diffusion coefficient and diffusion activation energy, and catalysis-related properties represented by acid strength and activation energy barrier. 2.4: The method of training multiple GNN models using Deep Ensembles and taking the variance, enabling Dropout for multiple sampling during MC Dropout inference, or directly outputting distribution parameters using Evidential Learning is used to quantify prediction uncertainty.
[0011] 3) Inject domain knowledge and tool invocation specifications into the large language model through SFT supervised fine-tuning. 3.1: Construct the SFT training dataset, which includes three types of data: domain knowledge question-answer pairs, tool call examples, and inference chain examples; 3.2: Adopt the structured dialogue format to standardize data, requiring all tool calls to include an inference block, strictly follow the JSON format specification, and provide natural language summaries after the tool returns results; 3.3: Through supervised fine-tuning, the model learns to master the professional terminology and conceptual system in the field of molecular sieves, learns to output Function Calls in accordance with the standard format, develops the reasoning habit of "thinking before acting", and understands the input and output formats of various calculation software.
[0012] 4) The GRPO algorithm combined with a multi-dimensional reward function is used to train the model through reinforcement learning. 4.1: Construct a set of reinforcement learning (RL) hints covering open-ended design tasks, constrained optimization tasks, and reverse reasoning tasks; 4.2: The GRPO algorithm is used to calculate the advantage through relative comparison within the group, without the need to train a Critic network; 4.3: Design a multi-dimensional instant reward function: 4.3.1: The weight of the format compliance reward is 0.15. Check whether the output contains a complete reasoning block, whether the JSON format of the tool call is valid, and whether the final answer contains the necessary elements. 4.3.2: The weight of the chemical logic reward is 0.25, which includes charge balance check, coordination rationality check, Pauling rule verification and topological rationality check; 4.3.3: The weight of the performance prediction reward is 0.60. The pre-trained GNN proxy model is called to make millisecond-level predictions. The prediction indicators include structural stability, target adsorption amount, and diffusion coefficient. The predicted values are normalized and used as the reward signal. 4.4: Perform the training process of sampling, evaluation, normalization, and updating. Generate multiple different output sequences for each prompt word. Calculate the comprehensive reward value of each sequence in parallel. Calculate the mean and standard deviation within the group to obtain the relative advantage. Use the proximal strategy to optimize the clip loss function of PPO and update the model parameters.
[0013] 5) Design a tool system based on the native Function Call capability of large language models, integrating fast prediction tools and high-precision calculation tools. 5.1: Design a fast prediction tool predict_property, which uses a GNN surrogate model to quickly predict the properties of molecular sieves. The parameters include the crystal structure in CIF format, the prediction type, and the conditional parameters. 5.2: Design a literature retrieval tool, search_literature, to retrieve literature, synthesis formulas, and experimental data related to molecular sieves. Parameters include search keywords, filtering conditions, and maximum number of results returned. 5.3: Design the structure generation tool generate_structure, which generates candidate molecular sieve structures based on constraints, including topology type, pore size range, silica-alumina ratio, target properties, and the number of candidates to generate; 5.4: Design a high-precision calculation tool, request_hpc_calculation, to submit high-precision calculation tasks to a high-performance computing (HPC) cluster for asynchronous execution. Parameters include the calculation method, CIF format structure, calculation parameters, and priority. 5.5: The trained model learns a "fast first, slow later" tool calling strategy, prioritizing the use of predict_property for quick filtering, and only triggering high-precision calculations after confirming the value.
[0014] 6) Construct an integrated architecture for a vertical intelligent system for molecular sieves, achieving asynchronous computation scheduling and multi-level caching optimization. 6.1: Build a user interface layer that supports multiple access methods including Web UI, API, CLI, and Jupyter Extension; 6.2: Construct the inference engine layer, adopt the large model inference framework vLLM to support continuous batch processing and paginated attention mechanism PagedAttention, and realize tool routing and context management; 6.3: Construct the service layer, including the GNN proxy model service deployed by the model service framework TorchServe, the knowledge retrieval system built with the full-text search engine Elasticsearch and the vector database, and the HPC computing cluster built with the high-performance computing scheduling system Slurm and the memory caching database Redis; 6.4: Implement an asynchronous scheduling mode for high-precision calculations. The system immediately returns the job ID and estimated time. The task is submitted to the queue for background execution, and the results are stored in the database upon completion. 6.5: Implement a multi-level caching mechanism, hash the input structure to generate a structure fingerprint, and return the cached result directly if the structure is the same.
[0015] Compared with the prior art, the present invention has the following beneficial technical effects and significant technical progress: 1) The vertical domain intelligent system architecture of the present invention enables the large language model to internalize molecular sieve expertise through deep domain training, and can realize four core functions: knowledge question answering, property prediction, reverse design and synthesis planning.
[0016] 2) This invention uses a GNN graph neural network surrogate model as the source of reward signals for reinforcement learning, which avoids calling time-consuming scientific computing software during training, greatly improves training efficiency, and achieves millisecond-level property prediction response.
[0017] 3) This invention proposes a "fast first, slow later" tool calling strategy, which prioritizes the use of the GNN proxy model for fast screening, and only triggers high-precision calculation after the value is confirmed, effectively balancing response speed and calculation accuracy.
[0018] 4) The multi-dimensional instant reward function design of this invention combines three dimensions: format compliance, chemical logic, and performance prediction, to ensure that the model output not only conforms to the standard format and meets the chemical principles, but also has good performance prediction capabilities. Attached Figure Description
[0019] Figure 1 This is a flowchart of the present invention; Figure 2 This is a schematic diagram illustrating the specific operation of Example 1. Detailed Implementation
[0020] To achieve intelligent question answering, property prediction, reverse design, and synthesis planning in the field of molecular sieves, this invention proposes a method for constructing a vertical domain intelligent system based on a large language model. This method first constructs a multi-source heterogeneous data processing flow to standardize structural data and extract literature information. Then, it injects domain knowledge and tool call specifications through supervised fine-tuning. Next, it uses the GRPO algorithm combined with a multi-dimensional reward function for reinforcement learning training. Simultaneously, it constructs a GNN proxy model to achieve millisecond-level property prediction, designs a tool system based on native Function Call capabilities, and finally constructs a system integration architecture to achieve asynchronous computation scheduling and multi-level caching optimization.
[0021] See appendix Figure 1 A method for constructing a molecular sieve vertical domain intelligent system based on a large language model, specifically including the following steps: 1) Construct a multi-source heterogeneous data processing workflow in the field of molecular sieves to achieve structural data standardization and literature information extraction. 1.1: Obtain official structure data for 250+ topological types from the International Zeolite Association (IZA) structure database, obtain 14,000+ computationally ready MOF structure data from the computationally ready metal-organic framework (CoRE) MOF database, obtain hypothetical structure data from the hypothetical crystal structure (PCOD) database, and obtain organic / organic metal crystal data from the Cambridge Crystal Structures (CSD) database. 1.2: Obtain molecular sieve-related papers from publishers such as Elsevier, Springer, and ACS; obtain preprints in the field of materials science from the academic preprint platform arXiv; and obtain molecular sieve synthesis patents from the USPTO, EPO, and CNIPA patent databases of the United States Patent and Trademark Office, the European Patent Office, and the China National Intellectual Property Administration. 1.3: Obtain material property data from DFT calculations from the Materials Genome Computational Database Materials Project, and computational materials science data from the Computational Materials Science Database NOMAD, to construct a giant canonical Monte Carlo (GCMC) and molecular dynamics (MD) simulation dataset; 1.4: All crystal structures are uniformly converted into three standardized formats: CIF format, SMILES-like topological encoding, and node-edge adjacency matrix graph representation; 1.5: Use natural language processing techniques to extract structured information such as synthetic formulations, property data, and experimental phenomena from literature.
[0022] 2) Construct a GNN (Graph Neural Network) surrogate model to achieve millisecond-level prediction and uncertainty quantification of molecular sieve properties. 2.1: The MEGNet graph neural network architecture is adopted as the basic architecture, and customized improvements are made for the characteristics of molecular sieves; 2.2: Construct structure-property pairs for training data, and train a GNN model to predict stability indices, adsorption properties, mass transfer properties, and catalysis-related properties; 2.3: Employ Deep Ensembles, MC Dropout, or Evidential Learning methods to quantify prediction uncertainty.
[0023] 3) Inject domain knowledge and tool invocation specifications into the large language model through SFT supervised fine-tuning. 3.1: Construct the SFT training dataset, which includes three types of data: domain knowledge question-answer pairs, tool call examples, and inference chain examples; 3.2: Adopt the structured dialogue format to standardize data, requiring all tool calls to include an inference block, strictly follow the JSON format specification, and provide natural language summaries after the tool returns results; 3.3: Through supervised fine-tuning, the model learns to master the professional terminology and conceptual system in the field of molecular sieves, learns to output Function Calls in accordance with the standard format, and develops the reasoning habit of "thinking before acting".
[0024] 4) The GRPO algorithm combined with a multi-dimensional reward function is used to train the model through reinforcement learning. 4.1: Construct a set of reinforcement learning (RL) hints covering open-ended design tasks, constrained optimization tasks, and reverse reasoning tasks; 4.2: The GRPO algorithm is used to calculate the advantage through relative comparison within the group; 4.3: Design a multi-dimensional instant reward function, including format compliance reward, chemical logic reward, and performance prediction reward; 4.4: Perform the training process of sampling, evaluation, normalization, and updating.
[0025] 5) Design a tool system based on the native Function Call capability of a large language model 5.1: Design a fast prediction tool, predict_property, and use a GNN surrogate model to quickly predict the properties of molecular sieves; 5.2: Design a literature retrieval tool, search_literature, to retrieve relevant literature and experimental data on molecular sieves; 5.3: Design the structure generation tool generate_structure to generate candidate molecular sieve structures based on constraints; 5.4: Design a high-precision calculation tool, request_hpc_calculation, to submit high-precision calculation tasks to a high-performance computing (HPC) cluster for asynchronous execution; 5.5: Train the model to learn a "fast first, slow later" tool calling strategy.
[0026] 6) Construct a system integration architecture to achieve asynchronous computing scheduling and multi-level caching optimization. 6.1: Construct a three-layer architecture consisting of a user interface layer, an inference engine layer, and a service layer; 6.2: Implementing an asynchronous scheduling mode for high-precision computing; 6.3: Implement a multi-level caching mechanism based on structural fingerprints.
[0027] The invention will be further described below with reference to specific examples and accompanying drawings.
[0028] Example 1 See Figure 2 This embodiment is a typical reverse design application scenario. Using the molecular sieve vertical field intelligent system constructed by this invention, intelligent recommendations from user needs to specific design solutions are realized. The specific process is as follows: 1) The user put forward the following requirement: "Please help me design a molecular sieve for separating propylene / propane, with a selectivity greater than 50".
[0029] 2) Analyze user-proposed requirements using a large language model to determine key constraints, including aperture and polarity; 3) Use the generate_structure tool to generate 5-10 candidate structures; 4) For each candidate, the predict_property tool is called for quick filtering. The GNN proxy model returns the predicted value and uncertainty within 100ms. 5) Select 2-3 optimal candidates and conduct detailed comparative analysis; 6) Provide the recommended structure and the rationale for the design; 7) Users can further request to call request_hpc_calculation to perform high-precision GCMC verification on the final solution.
[0030] Throughout the entire process described above, the large language model autonomously decides the timing and order of tool calls through its native Function Call capability, enabling intelligent recommendations from user needs to specific design solutions.
[0031] The above description is only a preferred embodiment of the present invention. Modifications may be made within the scope defined by the claims of the present invention, but all such modifications shall fall within the protection scope of the present invention.
Claims
1. A method for constructing a molecular sieve vertical domain intelligent system based on a large language model, characterized in that, The method includes the following steps: 1) Construct a multi-source heterogeneous data processing workflow in the field of molecular sieves to achieve structural data standardization and literature information extraction; 2) Construct a GNN graph neural network surrogate model to achieve millisecond-level prediction and uncertainty quantification of molecular sieve properties; 3) Inject domain knowledge and tool invocation specifications into the large language model through SFT-supervised fine-tuning; 4) The GRPO group relative policy optimization algorithm combined with a multi-dimensional reward function is used to train the model for reinforcement learning; 5) Design a tool system based on the native Function Call tool calling capability of large language models, integrating fast prediction tools and high-precision calculation tools; 6) Construct an integrated architecture for a vertical intelligent system for molecular sieves, and achieve asynchronous computation scheduling and multi-level caching optimization.
2. The method for constructing a vertical domain intelligent system for molecular sieves based on a large language model according to claim 1, characterized in that, Step 1) specifically includes: 1.1: Crystal structure data were obtained from the International Zeolite Association (IZA) structure database, the Computationally Ready Metal-Organic Frameworks (CoRE) MOF database, the Hypothetical Crystal Structures (PCOD) database, and the Cambridge Crystal Structures (CSD) database; literature data were obtained from academic literature and patent databases; and computational data were obtained from the Materials Project computational database and the NOMAD computational materials science database. 1.2: Convert all crystal structures into three standardized formats: CIF format, SMILES-like topological encoding, and node-edge adjacency matrix graph representation; 1.3: Use natural language processing techniques to extract structured information on synthetic formulations, property data, and experimental phenomena from literature.
3. The method for constructing a molecular sieve vertical domain intelligent system based on a large language model according to claim 1, characterized in that, Step 2) specifically includes: 2.1: Constructing a GNN proxy model based on the MEGNet architecture to analyze the characteristics of molecular sieves; 2.2: Extract computational data from the database, extract experimental measurement values from the paper, and run molecular simulation software RASPA and density functional theory DFT calculations on key structures to supplement the training set and construct structure-property pair training data; 2.3: Training a GNN surrogate model to predict stability indices, adsorption properties, mass transfer properties, and catalysis-related properties; 2.4: Employ Deep Ensemble, MC Dropout, or Evidential Learning methods to quantify prediction uncertainty.
4. The method for constructing a vertical domain intelligent system for molecular sieves based on a large language model according to claim 1, characterized in that, Step 3) specifically includes: 3.1: Construct an SFT training dataset that includes: domain knowledge question-answer pairs, tool call examples, and inference chain examples; 3.2: The text data and tool calls shall be standardized using the structured dialogue format. Each tool call shall include an inference block, follow the JSON format, and be summarized in natural language after the tool returns the results. 3.3: Through supervised fine-tuning, the large language model learns the professional terminology and conceptual system in the field of molecular sieves, learns to output Function Calls in accordance with the standard format, and develops the reasoning habit of "thinking before acting".
5. The method for constructing a molecular sieve vertical domain intelligent system based on a large language model according to claim 1, characterized in that, Step 4) specifically includes: 4.1: The GRPO algorithm is used for reinforcement learning training, and the advantage is calculated by comparing within the group. 4.2: The design includes a multi-dimensional instantaneous reward function encompassing format compliance rewards, chemical logic rewards, and performance prediction rewards; 4.3: Perform the training process of sampling, evaluation, normalization and updating, using a GNN surrogate model to provide reward signals.
6. The method for constructing a molecular sieve vertical domain intelligent system based on a large language model according to claim 5, characterized in that, Step 4.2 specifically includes: 4.2.1: Format compliance reward check whether the output contains a complete reasoning block, whether the tool call JSON format is valid, and whether the final answer contains the necessary elements; 4.2.2: Chemical logic rewards include: charge balance check, coordination rationality check, Pauling rule verification, and topological rationality check; 4.2.3: Performance Prediction Reward The pre-trained GNN proxy model is used for millisecond-level prediction. The prediction metrics include: structural stability, target adsorption amount and diffusion coefficient. The predicted values are normalized and used as the reward signal.