A method and system for determining a transition state of a multi-phase chemical reaction based on physical constraints and adaptive optimization, and a storage medium
By integrating nonlinear initial guessing, constrained optimal transport, uncertainty quantification, and adaptive optimization methods, this study solves the problems of nonlinear path fitting, physical realism, and adaptability to complex environments in the generation of chemical reaction transition states in existing technologies, and achieves efficient and reliable generation of multiphase chemical reaction transition states.
Patent Information
- Application Number
- CN202511750577.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-26
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2045-11-26
AI Technical Summary
Existing machine learning methods suffer from several problems when generating transition states for chemical reactions: insufficient ability to fit nonlinear paths, lack of physical realism constraints, limited application scenarios to isolated systems, inability to handle complex environmental factors, and an inability to balance efficiency and reliability.
A deterministic generation method for multiphase chemical reaction transition states based on physical constraints and adaptive optimization is adopted. By integrating a nonlinear initial guessing module, a constrained optimal transmission module, an uncertainty quantification module, and an adaptive optimization decision module, candidate transition state structures are generated and evaluated, and high-precision optimization is performed when necessary.
It achieves high-precision and robust simulation of complex chemical reactions, expands the application scenarios to multiphase and multicomponent chemical systems, balances computational efficiency and reliability, and is suitable for industrial-grade high-throughput screening.
Smart Images

Figure CN121215069B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of computational chemistry and artificial intelligence, in particular, to a method and system for determining transition state of multi-phase chemical reaction based on physical constraints and adaptive optimization, and a storage medium. The present application is particularly suitable for scenarios that require high precision and high efficiency to explore complex chemical reaction mechanisms, design new catalysts and functional materials. BACKGROUND
[0002] The transition state (TS) of a chemical reaction is a transient state with the highest energy and the most unstable structure that the system experiences in the process of transforming reactants into products. The structure and energy information of the transition state are of great theoretical significance for understanding the internal mechanism of chemical reactions, calculating reaction rates, predicting reaction paths, and designing new catalysts or functional materials that can efficiently catalyze specific reactions. Therefore, accurately and efficiently determining the transition state structure is one of the core problems in modern chemical and materials science research.
[0003] Traditional transition state search methods mainly rely on high-precision quantum chemical calculations, such as calculations based on density functional theory (DFT). These methods, such as the Climbing Image, Nudged Elastic Band (NEB) method and its variants, require complex point searching and optimization calculations on a multi-dimensional potential energy surface. Although these methods are theoretically reliable, their computational cost is extremely high, often requiring several hours or even days of CPU computation time to process a simple reaction. This huge computational overhead severely restricts the high-throughput screening of complex reaction networks and the large-scale virtual discovery of new materials, making the chemical research and development cycle long and costly.
[0004] In recent years, with the development of artificial intelligence technology, using machine learning to accelerate transition state search has become a promising research direction. Some machine learning-based generative models have been proposed to learn to generate transition state structures directly from given reactant and product structures, thereby reducing the computation time from hours to seconds or even shorter. However, existing machine learning methods still face a series of deep-seated technical problems in practical applications:
[0005] First, the fitting ability to the non-linear reaction path is insufficient. Many existing methods use the atomic coordinates of reactants and products for simple linear interpolation as the input of the initial guess of the generating model. This approach is acceptable for simple reactions with a reaction path close to linear, but for a large number of real reactions involving significant conformational changes, multi-atom rearrangements or complex cyclic structure changes, the real minimum energy path is highly nonlinear. There is a huge deviation between the simplified linear interpolation and the real path, which not only greatly increases the difficulty of the neural network model to learn the optimal path, but also may lead the model to converge to the wrong, high-energy non-optimal transition state structure, ultimately affecting the accuracy of the generation results.
[0006] Second, the model training process lacks effective constraints on physical reality. Many generating models, when training, usually only focus on minimizing the geometric deviation (such as root mean square error) between the predicted structure and the real structure. This purely data-driven training method makes the model like a "black box", which may learn some "mathematically optimal" but "physically unreasonable" atomic motion trajectories. For example, the model may generate structures with bond lengths or bond angles deviating significantly from the reasonable range, or the total energy of the system fluctuating dramatically during evolution. When faced with novel reaction types that do not appear in the training set, the generalization ability of such models and the physical reality of the generated structures will be severely challenged.
[0007] Third, the application scenarios are limited to idealized isolated systems. Most chemical reactions with practical value, such as industrial catalysis or enzyme reactions in living organisms, occur in complex environments, such as solution phase or on the surface of catalysts. The presence of solvent molecules or catalysts can significantly change the potential energy surface of the reaction, thereby deeply affecting the structure and energy of the transition state. However, most existing machine learning methods model the reactants and products as an isolated system in a vacuum, and cannot incorporate these key external environmental factors (such as catalysts, solvents) into the model, resulting in their inability to be applied to more extensive and more realistic multi-phase or multi-component chemical systems, which fundamentally hinders their application in more extensive and more realistic industrial-grade multi-phase or multi-component chemical systems, because in these scenarios, environmental effects often dominate the results of the reaction, leading to a greatly limited application range.
[0008] Fourth, the reliability and efficiency of the system cannot be balanced. Some advanced generative models are characterized by "single generation", which, although achieving very high computational efficiency, means that it provides a "once and for all, good or bad" result. When facing particularly difficult or at the edge of the model's ability to respond to cases, the quality of the structure of single generation may not be good, or even completely wrong. The system lacks an internal mechanism to assess the confidence of its generated results and adaptively decide whether to accept the results or start a more accurate but more costly optimization process. This mode makes the system sacrifice reliability while pursuing speed, and cannot achieve intelligent balance between efficiency and reliability. SUMMARY
[0009] The technical problem to be solved by the present application is to provide a multi-phase chemical reaction transition state deterministic generation method and system based on physical constraints and adaptive optimization, and a storage medium, which aims to not only maintain the high efficiency of machine learning methods, but also overcome the above problems and achieve accurate, robust and reliable simulation of complex and real chemical scenarios.
[0010] To achieve the above purpose, the present application provides a multi-phase chemical reaction transition state deterministic generation method based on physical constraints and adaptive optimization, which is executed by a computer system including a processor and a memory, and the method comprises the following steps:
[0011] Obtain initial structure data of reactants and products;
[0012] Call a nonlinear initial guess module, which performs path search based on a coarse-grained potential energy surface based on the initial structure data of the reactants and products, to generate a nonlinear initial guess path containing multiple intermediate mirror point structures;
[0013] Call a constraint optimal transport module, which takes the nonlinear initial guess path as input and performs path evolution in a pre-set constraint optimal transport model for multi-component systems to generate a candidate transition state TS structure; wherein the constraint optimal transport model imposes a pre-set geometric or energy constraint on environmental factors when dealing with environmental factors containing catalysts or solvents;
[0014] Call an uncertainty quantification module, which performs reliability evaluation on the candidate transition state structure and outputs a quantified confidence score;
[0015] Call an adaptive optimization decision module, which compares the confidence score with a pre-set confidence threshold and adaptively performs one of the following operations:
[0016] If the confidence score is higher than the high confidence threshold, the candidate transition state TS structure is directly accepted as the final result;
[0017] If the confidence score is lower than the low confidence threshold, a targeted optimization process is started, taking the candidate transition state TS structure as the initial point, to perform high-precision quantum chemical calculation to obtain a refined transition state TS structure;
[0018] The core of the present application is to propose and implement an intelligent decision-making closed loop with high degree of functional integration, and the hub is the constraint optimal transport module. This module combines optimal transport theory with chemical and physical constraints, not only solving the problem of how to efficiently generate transition states, but more importantly, it solves the difficult problem of how to accurately simulate in a complex and real environment containing catalysts or solvents. It organically connects the high-quality starting point provided by the nonlinear initial guess module and delivers physically highly reasonable candidate structures to the uncertainty quantification module and the adaptive optimization decision-making module, thereby ensuring the advanced nature, reliability and industrial applicability of the entire technical solution.
[0019] Further, based on the path search of the coarse-grained potential energy surface, specifically using the GFN2-xTB semi-empirical method or the classical force field method to quickly calculate the approximate minimum energy path between reactants and products, the reason for using this path as the input of the constraint optimal transport model is that the standard optimal transport can directly handle reactants and products, but for highly nonlinear atomic motion paths, pre-generated nonlinear initial paths can significantly reduce the evolution difficulty and ensure physical reality.
[0020] Further, when the environmental factor is a catalyst, the geometric constraint imposed on the environmental factor is specifically that during the path evolution process, the atomic positions of the catalyst are kept fixed or are limited to small vibrations near the pre-set active site.
[0021] Further, when the environmental factor is a solvent, the energy constraint imposed on the environmental factor is specifically that at each step of the path evolution, an implicit solvent model is used to perform real-time calculation and correction of the system's energy and interatomic forces.
[0022] Further, the uncertainty quantification module has a pre-trained graph neural network GNN model built in, which is used to predict the reliability of the generation of the candidate transition state TS structure.
[0023] Further, the targeted optimization process is specifically a geometry optimization calculation performed on the candidate transition state TS structure under the density functional theory DFT framework.
[0024] Further, a multi-phase chemical reaction transition state deterministic generation system based on physical constraints and adaptive optimization is also provided, which includes:
[0025] A processor;
[0026] A memory, in which computer program instructions executable by the processor are stored;
[0027] When the computer program instructions are executed by the processor, the system implements the following functional modules:
[0028] A nonlinear initial guess module, configured to perform a path search based on a coarse-grained potential energy surface, based on the obtained initial structure data of the reactants and products, to generate a nonlinear initial guess path containing a plurality of intermediate mirror point structures;
[0029] A constrained optimal transport module, configured to input the nonlinear initial guess path, and perform path evolution in a preset constrained optimal transport model for a multi-component system, to generate a candidate transition state TS structure; wherein the constrained optimal transport model imposes a preset geometric or energy constraint on environmental factors when processing environmental factors containing a catalyst or a solvent;
[0030] An uncertainty quantification module, configured to perform reliability evaluation on the candidate transition state TS structure, and output a quantified confidence score;
[0031] An adaptive optimization decision module, configured to compare the confidence score with a preset confidence threshold, and adaptively perform one of the following operations: if the confidence score is higher than a high confidence threshold, directly accepting the candidate transition state structure as a final result; and if the confidence score is lower than a low confidence threshold, starting a targeted optimization process, taking the candidate transition state TS structure as an initial point, and performing high-precision quantum chemical calculation to obtain a refined transition state TS structure.
[0032] Further, the constrained optimal transport module is configured to:
[0033] When the environmental factor is a catalyst, a position fixing constraint or a harmonic oscillator potential constraint is imposed on the atomic coordinates of the catalyst during path evolution, to realize geometric constraint;
[0034] When the environmental factor is a solvent, a polarizable continuum model PCM or a COSMO-like model is called to modify the Hamiltonian of the system at each step of path evolution, to realize energy constraint.
[0035] Further, the adaptive optimization decision module further includes a targeted optimization unit, which is configured to, when started by the adaptive optimization decision module, call quantum chemical calculation software in the memory, input coordinate data of the candidate transition state TS structure, and perform density functional theory DFT calculation.
[0036] Further, a computer readable storage medium is provided, and the computer program is stored on the computer readable storage medium, and when the computer program is executed by a processor, the computer program implements the above-mentioned method for determining a transition state of a multi-phase chemical reaction based on physical constraints and adaptive optimization.
[0037] The present application has the following advantages:
[0038] 1. Precision and robustness improvement: By adopting path-aware nonlinear initial guess, the present application provides a more reasonable and closer-to-real-physical-process starting point for subsequent accurate generation, reducing the learning difficulty of the neural network, thereby improving the transition state structure generation precision of complex and nonlinear reaction paths. In addition, by explicitly introducing physical constraints such as energy and bond energy in model training, the present application makes the model not only learn the geometric transformation between atoms, but also learn the atomic evolution process that conforms to the physical law, which enhances the model's generalization ability to new reaction types not seen in the training set, and ensures the physical reality of the generated structure, avoiding the generation of chemically unreasonable structures.
[0039] 2. Application range expansion: By constructing a constraint optimal transmission model for multi-component systems, the present application successfully expands the application scenario from ideal gas phase systems to more general and more practically meaningful catalytic reactions and solution phase reactions. The present application can incorporate catalysts or solvents as constrained environmental factors into the model, accurately simulating their influence on the reaction potential energy surface and transition state under the field effect, which is an expansion of the application field of existing isolated system models, enabling them to handle complex chemical systems in real industrial and scientific research scenarios.
[0040] 3. Balancing of system implementation efficiency and reliability: The present application constructs an intelligent decision-making closed loop integrating "fast generation-confidence evaluation-adaptive optimization". For the vast majority of reactions, the system can give high-confidence and reliable results at a relatively fast speed, achieving extremely high computational efficiency; while for a small number of difficult or marginal cases, benefiting from the much higher structure quality output by the constraint optimal transmission module than traditional methods, the system can automatically identify and invest a small amount of additional computing resources for "targeted treatment", and perform fine-tuning through high-precision methods, thereby avoiding the dilemma of excessive reliance on high-cost traditional methods or accepting low-quality results. This intelligent balancing mechanism supported by the high-quality output of the constraint optimal transmission module optimizes the overall performance and reliability of the system, making it have higher practical value in industrial-level high-throughput screening.
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description only constitute some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings. BRIEF DESCRIPTION OF DRAWINGS
[0042] Figure 1 Fig. 1 is a schematic diagram of the hardware and functional module architecture of the transition state certainty generation system according to an embodiment of the present application.
[0043] Figure 2 Fig. 2 is a flowchart of the transition state certainty generation method according to an embodiment of the present application.
[0044] Figure 3 Fig. 3 is a schematic diagram of the working principle of the nonlinear initial guess module according to an embodiment of the present application.
[0045] Figure 4 Fig. 4 is a schematic diagram of the loss function containing physical constraints used in the training phase of the constraint optimal transport model according to an embodiment of the present application.
[0046] Figure 5 Fig. 5 is a schematic diagram of the constraint optimal transport module processing a heterogeneous reaction system containing a catalyst according to an embodiment of the present application.
[0047] Figure 6 Fig. 6 is a schematic diagram of the working flow of the adaptive optimization decision module according to an embodiment of the present application. DETAILED DESCRIPTION
[0048] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments only constitute some embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0049] The present embodiment provides a highly integrated and intelligent, physical constraint and adaptive optimization based heterogeneous chemical reaction transition state certainty generation system and the corresponding implementation method. The system aims to generate physically real and chemically reasonable transition state structures with extremely high efficiency and accuracy, and can adapt to various complex chemical scenarios including gas phase, solution phase and catalytic reactions.
[0050] System hardware and software environment
[0051] Reference Figure 1The system of the present embodiment can be deployed on one or more high-performance computing servers. The server includes, but is not limited to, one or more central processing units (CPUs) 101, such as Intel Xeon or AMD EPYC series processors; one or more graphics processing units (GPUs) 103, such as NVIDIA A100 or H100, for accelerating neural network calculations; a large-capacity random access memory (RAM) 102, such as 256 GB or more, for storing program instructions, model parameters, and intermediate calculation data; a high-speed solid-state drive (SSD) or network storage 104 for persistent storage of chemical structure databases, trained model files, and calculation results; and a network interface 105 for data interaction with external databases or user terminals.
[0052] At the software level, a complete set of computer program instructions is stored in the memory 102, including an operating system (such as Linux), a deep learning framework (such as PyTorch or TensorFlow), a quantum chemistry calculation software package (such as Gaussian, ORCA), and the core functional module program of the present invention. When the processor 101 and the GPU 103 execute these instructions, the entire functionality of the present invention is realized.
[0053] System functional modules
[0054] As shown in Figure 1 , the system of the present invention is logically composed of four core functional modules, which are implemented by the processor 101 executing the corresponding program instructions in the memory 102:
[0055] Nonlinear initial guess module 10: responsible for generating an initial reaction path more accurate than traditional linear interpolation.
[0056] Constrained optimal transport module 20: responsible for evolving the initial path into a candidate transition state structure under physical and chemical constraints.
[0057] Uncertainty quantification (UQ) module 30: responsible for evaluating the reliability of the generated results.
[0058] Adaptive optimization decision module 40: responsible for intelligent decision-making of subsequent operations according to the reliability evaluation results.
[0059] The overall method flow will be described in detail below. Figure 2
[0060] Step S201: Obtain input data
[0061] At the beginning of the method, the system first acquires the input information of a specific chemical reaction. These information at least includes the initial three-dimensional structure data of the reactants (R) and products (P). These data are usually stored in standard chemical file formats (such as.xyz,.mol,.cif), which define the element type (such as C, H, O, N) of each atom and its coordinates (x, y, z) in the three-dimensional Cartesian coordinate system in detail. In some scenarios involving heterogeneous reactions, the input data may also include structural information of environmental factors (E), such as atomic coordinates of catalysts or types of solvents. These data are read by the processor 101 from the storage device 104 into the memory 102 for subsequent module processing.
[0062] Step S202: Generating path-aware nonlinear initial guess structure
[0063] This step is performed by the nonlinear initial guess module 10. Its core goal is to overcome the limitations of linear interpolation methods in the prior art, and the necessity of this module lies in overcoming the limitations of linear interpolation (such as Figure 3 as shown), providing a better starting point for constrained optimal transport.
[0064] Linear interpolation assumes that the atomic motion from reactants to products is along a straight line. However, the atomic motion path of real chemical reactions, especially those involving ring opening and closing, large-scale migration of groups, or complex rearrangement, is highly curved and nonlinear. Referring to Figure 3 , the linear interpolation path (dashed line) can be far from the real minimum energy path (solid line), and even can pass through a region with extremely high energy, which brings great and unnecessary difficulties to the subsequent precise model learning.
[0065] Therefore, this module adopts a "coarse-to-fine" strategy. It first performs a path search based on a coarse-grained potential energy surface with extremely low computational cost. Here, "coarse-grained potential energy surface" refers to an approximate energy description method with much less computational overhead than high-precision quantum chemical calculations. In a preferred embodiment, the path search based on the coarse-grained potential energy surface is specifically: using the GFN2-xTB semi-empirical method or the classical force field method to quickly calculate the approximate minimum energy path between the reactants (R) and the products (P).
[0066] GFN2-xTB is a semi-empirical tight-binding quantum chemistry method. Compared with expensive DFT calculations, it simplifies the calculation process through a large number of parameterizations, and can complete the energy and force calculations of medium-sized molecules in a few seconds, although the accuracy is lower, but it is enough to capture the general outline of the reaction path.
[0067] A classical force field is a faster approximation method that treats atoms as point masses connected by springs, using pre-parameterized functions (such as Lennard-Jones potential, bond length / bond angle / dihedral angle terms) to describe the interaction between atoms. It is extremely fast in calculation, but its scope of application and accuracy are limited by the force field parameters.
[0068] The processor 101 invokes the pre-installed GFN2-xTB program or force field calculation engine in the memory 102 to perform a simplified Nudged Elastic Band (NEB) calculation. The NEB method creates a series of intermediate structures "images" between the reactants and products, and moves these images through optimization algorithms (such as gradient descent) to eventually converge to the minimum energy path. Since GFN2-xTB is used, the entire process is usually completed within a few seconds.
[0069] Module output and data flow: The output of this module is no longer a single, linearly interpolated intermediate point, but a non-linear initial guess path composed of multiple (e.g. 8 to 16) intermediate image structures. This series of three-dimensional structure data is stored in a specific area of the memory 102 and passed to the next module.
[0070] The initial path generated by this method is more reflective of the topological features of the true reaction trajectory, such as the order of atomic motion, the conformation of key intermediates, etc., because it is optimized on an approximate but physically meaningful potential energy surface. This provides a high-quality starting point for the subsequent constrained optimal transport module 20, greatly reducing the difficulty of learning and optimization, and is the key first step to improve the accuracy of the final generated structure.
[0071] Step S203: Perform constrained optimal transport to generate candidate transition state
[0072] This step is performed by the constrained optimal transport module 20 and is the core calculation step of the invention; the module receives non-linear path input because it improves adaptability to complex environments, rather than direct processing of non-standard OT; the constrained optimal transport model is based on OT theory, but in the chemical context, pre-generated paths can inject physical prior knowledge and reduce bias for non-linear reactions, ensuring efficient evolution.
[0073] The module uses a pre-trained deep learning model based on the Optimal Transport (OT) theory.
[0074] Optimal Transport (OT) is a mathematical theory that aims to transform one probability distribution (or set of points) into another probability distribution (or another set of points) at the lowest cost. In this invention, it is used to find an "optimal" flow of atomic motion, "transporting" the atomic distribution on a nonlinear initial guess path to the actual, lowest-energy reaction path, and ultimately determining its highest-energy point—the transition state.
[0075] The constrained optimal transport module is crucial for applying this invention from a theoretical model to complex, real-world chemical scenarios. Its importance lies in the fact that it is not merely a simple structure generator, but rather deeply infused with physicochemical principles. Without this module, or by employing traditional generation models without physical constraints, the entire technical solution would degenerate. Firstly, it would be unable to handle real systems containing catalysts or solvents, severely limiting its application scope. Secondly, even in ideal systems, the generated structures might violate fundamental physical laws, leading to unreliable results. Therefore, this module, through its unique constraint mechanism, ensures the physical authenticity and broad applicability of the technical solution, serving as the cornerstone of this invention.
[0076] This solution differs from existing technologies; the OT model used in this module has been improved in key aspects:
[0077] Physical constraints were introduced during the training phase: such as Figure 4 As shown, the loss function L of this model during the training phase total It not only includes the traditional geometric deviation term L_OT (used to minimize the geometric difference between the predicted path and the true path), but also innovatively introduces a physical constraint regularization term λ*L. phys Here, λ is a hyperparameter used to balance the weights of the two terms. phys It can consist of one or more sub-items:
[0078] Energy gradient constraint: This penalty applies to atomic motions that cause drastic and unreasonable fluctuations in the total energy of the system during path evolution. It requires that the direction of atomic motion roughly follows the direction of the energy gradient (when far from the transition state), preventing the model from learning physically impossible energy abrupt changes.
[0079] • Chemical Bond Energy Constraints: This section utilizes chemical knowledge maps or prior knowledge to constrain the range of bond energies and bond lengths of chemical bonds that break and form during a reaction. For example, the breaking energy of a C / C single bond is typically within a certain range, and its bond length may elongate in the transition state, but it should not be stretched or shortened to an unreasonable degree indefinitely. This constraint penalizes structures that generate chemically unreasonable bond lengths or bond energies, ensuring the chemical rationality of the generated structures.
[0080] In this way, when performing gradient backpropagation and parameter updating of model training, the processor 101 not only considers geometric alignment, but also must "comply" with basic physical and chemical laws. This makes the trained model have physical insight, greatly improves the generalization ability and physical authenticity.
[0081] Model extension for multi-component systems: This module can handle multi-phase reaction systems containing environmental factors E (such as catalysts or solvents). It treats the entire system (R, P, E) as a multi-component system and imposes different constraints on different components; This feature enables the module to work organically with other parts, it receives high-quality multi-component system initial paths from the nonlinear initial guess module 10, generates a more physically and chemically reasonable candidate structure, provides a high-quality evaluation object for the subsequent uncertainty quantification module 30, and greatly improves the efficiency and reliability of the final decision of the adaptive optimization decision module 40.
[0082] In a preferred embodiment, when the environmental factor E is a catalyst, the geometric constraint imposed on the environmental factor E is specifically: keep the atomic positions of the catalyst fixed or limit them to small vibrations near the pre-set active site during path evolution. Refer to Figure 5 When the OT model processes the "transport" of atomic coordinates, the processor 101 will impose a position constraint on the atomic coordinates belonging to the catalyst C. This can be achieved in two ways: one is to fix its coordinates and not participate in optimization; The other is to impose a harmonic potential, that is, allow it to vibrate slightly near the equilibrium position to simulate real thermal vibration, but any large displacement will be subject to a huge energy penalty. In this way, the model can learn how the atoms of the reactants R and the products P move and recombine on the surface of the catalyst, while the surface of the catalyst itself remains stable.
[0083] In a preferred embodiment, when the environmental factor E is a solvent, the energy constraint imposed on the environmental factor E is specifically: during path evolution, use an implicit solvent model to calculate and correct the energy and interatomic forces of the system in real time.
[0084] The implicit solvent model, such as the Polarizable Continuum Model (PCM) or the COSMO-like model, is a method that approximates the solvent environment as a continuous medium with a specific dielectric constant. It avoids explicitly handling thousands of solvent molecules, thereby greatly saving computational resources.
[0085] At each step of the path evolution, when the OT model needs to calculate the energy and interatomic forces of the current system, the processor 101 calls a PCM calculation module. This module calculates the additional energy and force contributions brought by the solvation effect according to the current atomic configuration and the dielectric constant of the solvent, and returns these corrected energy and force to the OT model. This means that the entire path evolution process is carried out in a simulated solution environment, so that the stabilizing (or destabilizing) effect of the solvent on the transition state structure and energy can be accurately captured.
[0086] It is through the creative combination of the above physical constraints and multi-component modeling that this module ensures that the candidate transition state structure output is not only geometrically accurate, but also highly reasonable in physics and chemistry, and can truly reflect the influence of the complex environment. The logic of ensuring high-quality output from the source makes this solution achieve high precision, high reliability and wide applicability.
[0087] Module output and data flow: After the execution of the constraint optimal transmission module 20, a single, high-precision candidate transition state TS structure is output. The three-dimensional atomic coordinate data of this structure is stored in the memory 102, ready for reliability evaluation.
[0088] This module solves the core problem of poor physical reality and limited application scenarios of the prior art through physical constraints and multi-component modeling. The candidate TS structure generated by it is not only geometrically accurate, but also more reasonable in physics and chemistry, and can reflect the influence of the real reaction environment.
[0089] Step S204: Perform uncertainty quantification evaluation
[0090] This step is performed by the uncertainty quantification UQ module 30. Its purpose is to provide a basis for subsequent intelligent decision-making.
[0091] Although the candidate TS generated in step S203 is of high quality, for some extremely difficult or model insufficiently learned reactions, its result may still have uncertainty. The role of the UQ module 30 is to quantify this uncertainty.
[0092] In a preferred embodiment, the uncertainty quantification module 30 is built-in with a pre-trained graph neural network GNN model for predicting the generation reliability of the candidate transition state TS structure.
[0093] Graph neural network GNN is a deep learning model specially designed for processing graph-structured data. In this invention, the molecular structure is naturally regarded as a graph, where atoms are nodes and chemical bonds are edges.
[0094] The GNN model is trained for a classification task: input a reaction system (including reactants R, products P, and a candidate TS generated by module 20), output the probability that the candidate TS is "reliable". This probability value (between 0 and 1) is used as the confidence score.
[0095] The processor 101 inputs the structural data (atomic types, coordinates, bond connection relationships) of R, P, and the candidate TS into this pre-trained GNN model. The GPU 103 is responsible for performing the forward propagation calculation of the GNN, and obtains a confidence score in a very short time (usually milliseconds).
[0096] Module output and data flow: The UQ module 30 outputs a floating-point number, i.e. the confidence score. The score is passed to the adaptive optimization decision module 40 together with the candidate TS structure.
[0097] This module adds a "self-reflection" capability to the entire system. It makes the system no longer blindly output results, but can give a reliability judgment on its own output, which is the basis for intelligent decision-making.
[0098] Step S205: Perform adaptive optimization decision
[0099] This step is performed by the adaptive optimization decision module 40, which is the "command center" of the present application to achieve the intelligent balance between efficiency and reliability.
[0100] Reference Figure 6 The logic of this module is very clear. It receives the confidence score from module 30 and compares it with two preset threshold values: a high confidence threshold (e.g. 0.9) and a low confidence threshold (e.g. 0.7).
[0101] High confidence path (e.g. score > 0.9): If the confidence score is very high, it indicates that the UQ module is very confident about the quality of the candidate TS. In this case, module 40 decides to accept the candidate transition state TS structure directly as the final result. The system saves the structure data to storage device 104 or outputs it to the user. This path achieves the highest computational efficiency and is suitable for most "simple" reactions.
[0102] Low confidence path (e.g. score < 0.7): If the confidence score is very low, it indicates that the UQ module considers the quality of the candidate TS to be unreliable. In this case, module 40 decides to start the targeted optimization process.
[0103] In a preferred embodiment, the targeted optimization process is specifically: performing a geometry optimization calculation on the candidate transition state (TS) structure under the framework of density functional theory (DFT).
[0104] Density functional theory (DFT) is one of the most widely used high-precision quantum chemistry calculation methods. It can handle more complex systems than traditional ab initio methods while maintaining high accuracy.
[0105] Module 40 will use this low-confidence candidate TS structure as the initial point for a traditional but high-precision DFT transition state optimization calculation. This is crucial because the candidate TS generated by module 20, even if it has low confidence, is much better structured than random guesses or linear interpolation results. Using such a high-quality initial point can greatly speed up the convergence of DFT optimization calculations, which originally took hours to complete may be completed in a few tens of minutes. Processor 101 calls quantum chemistry software package (such as Gaussian) in memory 102 to perform this calculation and obtain a refined, high-confidence transition state TS structure as the final result.
[0106] Medium confidence path (0.7 <= score <= 0.9): For structures with confidence in the middle interval, the system can handle according to the preset strategy, for example, directly accept to balance efficiency, or also start targeted optimization to ensure the highest quality.
[0107] Module output and data flow: The module finally outputs a determined, high-quality transition state structure.
[0108] Adaptive optimization decision module 40 perfectly combines speed and accuracy. It acts like an experienced expert, quickly handling simple tasks and investing extra effort on difficult tasks, thereby minimizing the average computational cost while ensuring the reliability of the results.
[0109] Above, through the coordinated work of the four modules, a complete, closed-loop, intelligent solution is formed. Nonlinear initial guess module 10 provides a high-quality starting point for the entire process; constrained optimal transport module 20 uses physical constraints and multi-component modeling capabilities to generate more physically and chemically reasonable candidate structures; uncertainty quantification module 30 plays the role of "quality inspector"; adaptive optimization decision module 40 intelligently selects the most efficient path based on the test results. The four modules are closely linked and support each other in function, working together to solve the technical problems of existing technology in terms of accuracy, physical realism, application range, and reliability, and ultimately achieve significant value in the field of computational chemistry. The entire scheme combines algorithm features (OT, GNN) with technical features (chemical structure data processing, physical law constraints, quantum chemistry calculation calling) closely, its input is the data representation of physical entities, its output is the accurate prediction of the state of physical entities, and its process is a great improvement in computer internal performance (such as calculation speed, efficiency, reliability).
[0110] The above merely provides the preferred embodiments of the application, and not intended to limit the protection scope of the application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the application shall be included in the protection scope of the application.
Claims
1. A deterministic method for generating transition states of multiphase chemical reactions based on physical constraints and adaptive optimization, characterized in that, The method is executed by a computer system, the computer system including a processor (101) and a memory (102), and the method includes the following steps: Obtain initial structural data of reactants (R) and products (P); The nonlinear initial guessing module (10) is invoked. Based on the initial structure data of the reactant (R) and product (P), the nonlinear initial guessing module (10) performs a path search based on the coarse-grained potential energy surface to generate a nonlinear initial guessing path containing multiple intermediate mirror point structures. The constrained optimal transport module (20) is invoked. The constrained optimal transport module (20) takes the nonlinear initial guess path as input and performs path evolution in a preset constrained optimal transport model for multi-component systems to generate candidate transition state TS structures. The constrained optimal transport model applies preset geometric or energy constraints to the environmental factors (E) containing catalysts or solvents when dealing with them. The uncertainty quantization module (30) is invoked to perform a reliability assessment on the candidate transition state (TS) structure and output a quantified confidence score. The adaptive optimization decision module (40) is invoked, which compares the confidence score with a preset confidence threshold and adaptively performs one of the following operations: If the confidence score is higher than the high confidence threshold, the candidate transition state (TS) structure is directly accepted as the final result; If the confidence score is lower than the low confidence threshold, the targeted optimization process is initiated, and the candidate transition state (TS) structure is used as the starting point to perform high-precision quantum chemical calculations to obtain the refined transition state (TS) structure.
2. The deterministic generation method for multiphase chemical reaction transition states based on physical constraints and adaptive optimization according to claim 1, characterized in that, The path search based on the coarse-grained potential energy surface specifically employs the GFN2-xTB semi-empirical method or the classical force field method to quickly calculate the approximate minimum energy path between the reactant (R) and the product (P).
3. The deterministic generation method for multiphase chemical reaction transition states based on physical constraints and adaptive optimization according to claim 1, characterized in that, When the environmental factor (E) is a catalyst, the geometric constraint applied to the environmental factor (E) is specifically: during the path evolution process, the atomic position of the catalyst is kept fixed, or it is restricted to vibrate slightly near a preset active site.
4. The deterministic generation method for multiphase chemical reaction transition states based on physical constraints and adaptive optimization according to claim 1, characterized in that, When the environmental factor (E) is a solvent, the energy constraint imposed on the environmental factor (E) is specifically as follows: at each step of the path evolution, the implicit solvent model is used to calculate and correct the energy and interatomic forces of the system in real time.
5. The deterministic generation method for multiphase chemical reaction transition states based on physical constraints and adaptive optimization according to claim 1, characterized in that, The uncertainty quantification module (30) incorporates a pre-trained graph neural network (GNN) model to predict the generation reliability of the candidate transition state (TS) structure.
6. The deterministic generation method for multiphase chemical reaction transition states based on physical constraints and adaptive optimization according to claim 1, characterized in that, The targeted optimization process specifically involves performing geometric optimization calculations on the candidate transition state (TS) structure within the density functional theory (DFT) framework.
7. A deterministic generation system for the transition state of a multiphase chemical reaction based on physical constraints and adaptive optimization, characterized in that, The system includes: Processor (101); A memory (102) storing computer program instructions that can be executed by the processor (101); When the computer program instructions are executed by the processor (101), the system implements the following functional modules: The nonlinear initial guessing module (10) is used to perform a path search based on the coarse-grained potential energy surface based on the initial structure data of the reactants (R) and products (P) to generate a nonlinear initial guessing path containing multiple intermediate mirror point structures. The constrained optimal transport module (20) is used to take the nonlinear initial guess path as input and perform path evolution in a preset constrained optimal transport model for multi-component systems to generate candidate transition state TS structures; wherein, when the constrained optimal transport model deals with environmental factors (E) containing catalysts or solvents, it imposes preset geometric or energy constraints on the environmental factors (E); Uncertainty quantification module (30) is used to perform reliability assessment on the candidate transition state TS structure and output a quantified confidence score; The adaptive optimization decision module (40) is used to adaptively perform one of the following operations based on the comparison between the confidence score and a preset confidence threshold: if the confidence score is higher than the high confidence threshold, the candidate transition state (TS) structure is directly accepted as the final result; and if the confidence score is lower than the low confidence threshold, a targeted optimization process is initiated, and the candidate transition state TS structure is used as the starting point to perform high-precision quantum chemical calculations to obtain a refined transition state TS structure.
8. The deterministic generation system for multiphase chemical reaction transition states based on physical constraints and adaptive optimization according to claim 7, characterized in that, The constrained optimal transmission module (20) is configured as follows: When the environmental factor (E) is a catalyst, the geometric constraint is achieved by applying position-fixed constraints or harmonic oscillator potential constraints to the atomic coordinates of the catalyst during the path evolution process. When the environmental factor (E) is a solvent, the energy constraint is achieved by invoking a polarized continuum model (PCM) or a COSMO-like model to correct the Hamiltonian of the system at each step of the path evolution.
9. The deterministic generation system for multiphase chemical reaction transition states based on physical constraints and adaptive optimization according to claim 7, characterized in that, The adaptive optimization decision module (40) further includes a targeted optimization unit, which is configured to: when started by the adaptive optimization decision module (40), call the quantum chemical calculation software in the memory (102), take the coordinate data of the candidate transition state TS structure as input, and perform density functional theory DFT calculation.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements a deterministic generation method for multiphase chemical reaction transition states based on physical constraints and adaptive optimization as described in any one of claims 1 to 6.
Citation Information
Patent Citations
System and method for improved computer drug design
US20060041414A1
Determining the intrinsic reaction coordinate of a chemical reaction by nested path integrals
US20240321408A1