Quasi-reaction path adaptive enhanced sampling method

By combining semi-empirical quantum chemistry methods and active learning strategies, high-quality reaction machine learning potential function training data is generated, solving the problem of high computational cost in complex catalytic systems. This achieves efficient and low-cost data generation and model training, improving the accuracy and success rate of the model.

CN122067658APending Publication Date: 2026-05-19DALIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610221752.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-25
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing technologies require a large amount of computational resources when constructing high-precision reaction machine learning potential functions, especially in complex catalytic systems. Directly using high-precision DFT methods for transition state search and data labeling is computationally expensive and inefficient, making it difficult to generate high-quality training data.

Method used

The semi-empirical quantum chemistry method GFN2-xTB is used for transition state search and intrinsic reaction coordinate path calculation. Combined with normal mode sampling and active learning strategies, a high-quality training dataset is generated. Random perturbations are introduced through normal mode sampling and data is filtered using a committee query strategy, which reduces computational cost and improves data coverage.

Benefits of technology

It significantly reduced the computational cost and time for data generation, improved the accuracy and efficiency of the model, and achieved the construction of a high-precision reactive machine learning potential function. The RMSE of transition state structure optimization was reduced from 0.751 Å to 0.307 Å, the computation time was reduced from 310 hours to 39 hours, and the success rate reached 93%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122067658A_ABST
    Figure CN122067658A_ABST
Patent Text Reader

Abstract

The invention provides a quasi-reaction path adaptive enhanced sampling method, which is used for constructing a reaction machine learning potential database. The method comprises the following steps: firstly, generating a catalyst ligand candidate space by using a computer-aided molecular design tool based on a skeleton, and quickly generating a transition state initial conjecture through a GENiniTS-RS algorithm; the method comprises the core steps of combining an intrinsic reaction coordinate path of a semi-empirical quantum chemistry method (GFN2-xTB), and performing disturbance sampling around an IRC path by using normal modulus sampling so as to capture complex characteristics of a reaction potential energy surface. Then, a committee queries an active learning strategy to screen a high-information-entropy configuration to carry out high-precision density functional theory (DFT) marking; compared with full DFT sampling, the method has the advantages that the calculation cost is remarkably reduced, meanwhile, the success rate and precision of a machine learning model in transition state optimization are improved through enhanced sampling, and key data support is provided for high-throughput screening of efficient catalysts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of chemical product engineering and computer-aided molecular design, and crosses catalytic reaction engineering, artificial intelligence and data-driven modeling. Specifically, it involves the development of a quasi-reaction path adaptive enhancement sampling method for the potential energy surface fitting needs of complex chemical reactions (such as organometallic catalysis processes), which is used to efficiently prepare training data for reaction machine learning potential functions. Background Technology

[0002] The search for transition states (TS) and the screening of organic ligands are crucial in designing highly active organometallic catalysts with efficient reaction kinetics. However, accurate identification of transition states often requires a large amount of computationally intensive quantum chemical calculations, such as density functional theory (DFT).

[0003] In recent years, reactive machine learning potentials (RMLPs) have been able to reduce computational complexity by learning the relationship between the three-dimensional structure of molecules and the potential energy surface (PES). O ( N 3-4 Reduced to O ( N High-precision RMLP models have become an effective alternative to traditional DFT. However, constructing high-precision RMLP models requires a large number of continuous data points along the intrinsic reaction coordinate (IRC) path. Directly using high-precision DFT methods for conformational search and data labeling along the IRC path consumes massive amounts of computational resources (e.g., full DFT sampling for a single reaction step may require hundreds of thousands of core hours), which has become the biggest bottleneck restricting the application of machine learning potential functions in complex catalytic systems. Therefore, how to generate datasets containing high-quality transition states and IRC path conformations in a low-computational-cost and automated manner, while ensuring that the data can fully describe the complex potential surface near the reaction coordinate, has become a key technical challenge that urgently needs to be addressed. Summary of the Invention

[0004] This invention aims to provide a quasi-reaction path adaptive enhancement sampling method that is computationally inefficient, highly automated, and has broad sampling coverage. By integrating semi-empirical quantum chemistry methods, normal mode sampling (NMS), and active learning filtering strategies, it significantly reduces the computational cost during data acquisition and provides high-quality training data for machine learning potential functions (such as models based on the MACE architecture).

[0005] The method of the present invention includes the following specific steps: Step 1: Automatic assembly of catalyst ligands based on molecular framework and initial generation of transition state Input pre-defined reactants and ligand fragments, and use computer-aided molecular design (CAMD) tools to generate a catalyst library represented by SMILES strings; then, use the transition state preliminary structure generation algorithm (GENiniTS-RS) based on SMARTS-encoded reaction templates and pre-calculated three-dimensional coordinates of reaction sites to automatically convert them into preliminary three-dimensional structures of transition states.

[0006] Step 2: Rapid search for low-cost quasi-reaction pathways To address the initial guesses about the transition states, a low-cost semi-empirical quantum chemical method (GFN2-xTB) was employed to rapidly search for transition states and calculate intrinsic reaction coordinates (IRC) paths. This yielded preliminary geometric structures of key conformations for the reaction pathways, and frequency analysis calculations were performed.

[0007] Step 3: Adaptive Enhanced Sampling along Quasi-Reaction Path To enhance the machine learning model's ability to perceive local potential energy surfaces, based on the normal vibrational modes extracted in step 2... ) and their corresponding force constants ( The configuration along the IRC path is randomly perturbed and sampled. The specific displacement calculation formula is as follows: in, For along the first Coordinate perturbation of each vibration mode For random symbols, A pseudo-random number between [0,1] Number of atoms For sampling temperature, Let be the Boltzmann constant. The formula for calculating the Cartesian coordinates of the new conformation after perturbation is: Step 4: Intelligent Data Filtering Based on Active Learning The massive number of conformations generated in step 3 are filtered using a Query-by-Committee (QbC) active learning strategy. Multiple committee models with different initialization conditions are constructed to predict on the candidate dataset, using the standard deviation of atomic forces as the divergence metric. In each iteration, redundant data is removed, and only conformation data with large model prediction discrepancies and high bias are retained and added to the core training set.

[0008] Step 5: High-precision data annotation and model training High-level density functional theory (e.g., ...) is used only on the core training set selected in step 4. b97xd / def2svp performs precise labeling of energy and forces, and inputs it into a high-order equivariant graph neural network (such as the MACE architecture) for training a reactive machine learning potential (RMLP) model.

[0009] The beneficial effects of this invention are as follows: By integrating semi-empirical quantum chemistry methods, normal mode sampling (NMS), and active learning strategies, this invention significantly reduces the cost of preparing training data for reaction machine learning potential functions, achieving a dual breakthrough in computational efficiency and model accuracy. Specifically, in the data generation stage, GFN2-xTB is used to replace traditional DFT for transition state search and IRC path calculation, improving the efficiency of the underlying path search by nearly four orders of magnitude. In the data augmentation stage, random perturbations are introduced near the reaction coordinates through normal mode sampling, effectively solving the problem of model collapse at the saddle point (transition state) of the potential energy surface, reducing the RMSE of transition state structure optimization from 0.751 Å to 0.307 Å. In the data labeling stage, the committee query (QbC) active learning strategy is used to intelligently filter high information gain samples, reducing the DFT computation time from 310 hours to 39 hours (a reduction of 87.4%), while ensuring that the model achieves R-values ​​in energy and atomic force predictions. 2 With a high accuracy of approximately 1.000 and a transition state search success rate of up to 93% on the external test set, this method provides an automated and low-cost complete solution for the efficient construction of machine learning potential functions in complex catalytic systems. Attached Figure Description

[0010] Figure 1This document describes a hybrid database generation strategy for constructing a transition state dataset for the rhodium-catalyzed ethylene hydrogenation reaction, which involves multiple phosphorus ligands. a) represents the ethylene hydrogenation reaction and its pre-calculated 3D atomic coordinates of the reaction sites; b) uses the framework-based CAMD tool to generate organometallic catalysts represented by SMILES containing various phosphorus ligands; c) uses GENiniTS-RS to generate initial guesses of the transition states; d) uses GFN2-xTB for conformational sampling and DFT for labeling molecular energies and atomic forces; e) describes an active learning strategy for filtering the dataset.

[0011] Figure 2 This is a comparison of the chemical feature space and atomic force distribution. a) shows the chemical space mapped by transition state using the ACSF descriptor and t-SNE method, along with the molecular conformation along the IRC path and the NMS conformation (green dots represent transition states (TS), blue dots represent molecular conformations along the IRC path, and gray dots represent NMS conformations); b) shows the same chemical space as a, displaying the distribution of newly added data points during active learning (gray dots represent the complete dataset, green dots represent the initial data points from active learning training, and red dots represent newly added data points through the active learning strategy); c) is a violin plot of atomic force distribution considering both IRC and NMS conformations; d) is a violin plot of atomic force distribution considering only the IRC conformation.

[0012] Figure 3 This is a flowchart of the method of the present invention. Detailed Implementation

[0013] To better understand this invention patent, the implementation steps and beneficial effects of this invention will be further explained in detail below with reference to a specific application example (ligand screening for ethylene hydrogenation reaction catalyzed by rhodium-based catalysts).

[0014] A quasi-reaction path adaptive enhancement sampling method includes the following specific steps: Step 1: Catalyst Assembly and Initial Geometry Generation. A SMILES representation of the catalyst is generated using a mathematical programming-based molecular design tool, and then converted into an initial conjectured transition state configuration.

[0015] Step 2: Rapid exploration of quasi-reaction pathways. Low-cost semi-empirical quantum chemistry methods are used to search for transition states and sample IRC pathways from the initial conjecture to obtain preliminary reaction trajectories.

[0016] Step 3: Adaptive Enhanced Sampling. Based on the IRC path configuration, random perturbations are introduced near the potential energy surface using Normal Mode Sampling (NMS) to generate an enhanced sampling dataset.

[0017] Step 4: Data Filtering and High-Precision Labeling. An active learning strategy is adopted, using the standard deviation of predictions from multiple committee models to select representative configurations, and density functional theory (DFT) is used to calculate their energies and atomic forces to form the final training set.

[0018] In step 1, the initial conjecture of the transition state is generated using the GENiniTS-RS algorithm, which automatically converts SMILES into a three-dimensional initial transition state structure based on the SMARTS-encoded reaction template and the three-dimensional coordinates of the reaction site.

[0019] Normal mode sampling (NMS) in step 3 extracts the orthogonal vibrational modes and force constants of molecules and applies random amplitude displacements to the atomic coordinates, thereby enriching the model's understanding of the potential energy surface near the reaction coordinates and improving the model's robustness in transition state optimization.

[0020] The active learning strategy in step 4 employs the query-by-committee (QbC) method, which identifies configurations that are difficult for the model to learn by evaluating the prediction bias of atomic forces, thereby minimizing the cost of data labeling.

[0021] Example: Preparation of mixed data and potential function fitting for rhodium (Rh)-catalyzed ethylene hydrogenation reaction 1. Initial guess: batch generation of conformations First, a mixed-integer linear programming (MILP) model combined with the BARON solver was used to automatically assemble and generate 2073 Rh-centered organometallic catalysts containing different phosphorus ligands. The GENiniTS-RS algorithm was then used to rapidly transform the reactants and catalysts into a preliminary three-dimensional transition state structure.

[0022] 2. Semi-empirical IRC path search and NMS data augmentation 1069 catalysts were randomly selected, and the semi-empirical method GFN2-xTB was used for transition state search and IRC calculation. To enrich the data diversity, 10 NMS perturbation samples were generated for each IRC node.

[0023] Comparison of beneficial effects: Taking a typical reaction system containing 44 atoms as an example, on a 28-core Intel Xeon CPU, the total time taken to complete transition state, IRC, frequency analysis, and NMS using GFN2-xTB is only about 392.7 cores; if the traditional DFT method is used to complete the same task, it would consume about 3,936,643 cores. The underlying path search of this invention improves efficiency by nearly four orders of magnitude.

[0024] 3. Active learning filtering and precise DFT annotation The above steps initially generated a candidate pool of 496,690 molecular conformations. If all were labeled using DFT, it would take 310 hours on an 80-node high-performance cluster. This invention introduces a QbC active learning mechanism, which sets a threshold based on the force prediction variance of the committee model. After five rounds of active learning iterations, highly redundant data is automatically filtered out, ultimately selecting only 62,545 of the most informative and high-biased data points. DFT calculation and labeling of these points drastically reduced the total time to 39 hours.

[0025] 4. Machine Learning Model Training and Validation Metrics The model was trained using the MACE network architecture on 62,545 cleaned and labeled datasets. The model's performance on the external test set is as follows:

[0026] As shown in the table above, the model trained using the enhanced sampling and active learning algorithm of this invention (AL MACE w / NMS) achieves prediction accuracy close to the high-precision DFT benchmark, with a root mean square error (RMSE) of only 0.307 Å for the transition state geometry on the test set. Furthermore, the model achieves a transition state search success rate of up to 93% across 100 external test responses. Compared to models trained directly using semi-empirical methods or training sets without NMS perturbation, this invention not only completely solves the problem of model collapse at the saddle point (transition state) on the potential energy surface but also significantly reduces the computation time and resource consumption required for initial data preparation.

Claims

1. A quasi-reaction path adaptive enhancement sampling method, characterized in that, The specific steps include the following: Step 1: Catalyst assembly and initial geometry generation; using mathematical programming-based molecular design tools to generate the SMILES representation of the catalyst and convert it into the initial conjectured transition state configuration; Step 2: Rapid exploration of quasi-reaction pathways; using low-cost semi-empirical quantum chemistry methods to search for transition states and sample IRC pathways in the initial hypothesis to obtain preliminary reaction trajectories; Step 3: Adaptive Enhanced Sampling; Based on the IRC path configuration, random perturbations are introduced near the potential energy surface using normal mode sampling to generate an enhanced sampling dataset; Step 4: Data filtering and high-precision labeling; An active learning strategy is adopted, using the standard deviation of multiple committee models to select representative configurations, and density functional theory is used to calculate their energy and atomic forces to form the final training set.

2. The quasi-reaction path adaptive enhancement sampling method according to claim 1, characterized in that, In step 1, the initial conjecture of the transition state is generated using the GENiniTS-RS algorithm, which automatically converts SMILES into a three-dimensional initial transition state structure based on the SMARTS-encoded reaction template and the three-dimensional coordinates of the reaction site.

3. The quasi-reaction path adaptive enhancement sampling method according to claim 1, characterized in that, In step 3, normal mode sampling extracts the orthogonal vibrational modes and force constants of molecules and applies random amplitude displacements to the atomic coordinates, thereby enriching the model's understanding of the potential energy surface near the reaction coordinates and improving the model's robustness in transition state optimization.

4. The quasi-reaction path adaptive enhancement sampling method according to claim 1, characterized in that, The active learning strategy in step 4 employs a committee query method to identify configurations that are difficult for the model to learn by evaluating the prediction bias of atomic forces, thereby minimizing the cost of data labeling.