Multi-scale heuristic chemical reaction network automatic exploration method and system

Through the multi-scale heuristic chemical reaction network automatic exploration system, the problem of low search efficiency of complex chemical reaction paths in the existing technology is solved, efficient and automated reaction path exploration and verification are achieved, and the reliability of the calculation results is improved.

CN120183545APending Publication Date: 2025-06-20NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510272957.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

When facing complex and diverse chemical reactions, the prior art has low computational efficiency, incomplete paths and poor objectivity, making it difficult to achieve automated and efficient reaction path search.

Method used

The multi-scale heuristic chemical reaction network automatic exploration system is used to perform molecular conformation pre-processing through a molecular structure processor, and the determination rules inspired by chemical knowledge are combined to automatically identify high-active sites and pre-exclusion of inefficient paths. The reaction path is explored using hierarchical screening and gradual optimization strategies, and path validation and data analysis are carried out.

Benefits of technology

It realizes efficient search, automatic verification and data analysis of reaction paths, improves the reliability and calculation efficiency of results, and reduces manual intervention and invalid calculations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183545A_ABST
    Figure CN120183545A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-scale heuristic chemical reaction network automatic exploration method and system. The method comprises the following steps: S1, carrying out molecular conformation pretreatment on a plurality of molecular structures participating in a chemical reaction; s2, in combination with a judgment rule inspired by chemical knowledge, performing automatic identification of high-activity sites and pre-elimination of low-efficiency paths, and reserving candidate reaction site pairs with polarity matching and reasonable space proximity by utilizing hierarchical screening; s3, performing high-precision path optimization on the preliminary reaction path hypothesis by adopting a hierarchical screening-progressive optimization calculation strategy so as to explore a reaction path; s4, obtaining a reliable primitive reaction path, and constructing a complete reaction network and a convergence judgment mechanism based on a network expander; and S5, carrying out visual output through an interactive analysis interface. According to the method, efficient search, automatic verification and data analysis of the reaction path are realized, and the result reliability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computational chemistry and reaction path prediction, and particularly relates to a multi-scale heuristic automatic exploration system for chemical reaction networks and its system. Background Art

[0002] In computational chemistry and catalytic research, reaction path search is a key step in determining chemical reaction mechanisms. Traditional reaction path exploration methods mainly rely on manual design and manual optimization, such as the local transition state (TS) search method, in which researchers usually use the eigenvector following method to optimize the saddle point and calculate the intrinsic reaction coordinate (IRC) to connect the TS and adjacent intermediates, as Figure 1 shown. Traditional manual calculation of reaction paths usually relies on the researcher's experience and intuition about reaction mechanisms, and often requires step-by-step optimization of the path from reactants to products. Although this method can achieve good results in some simple reaction systems, its disadvantages and limitations are also very obvious, such as high dependence on experience and guesswork, low computational efficiency, incomplete path exploration, strong subjectivity, etc. Therefore, traditional manual calculation of reaction paths often has problems such as low computational efficiency, incomplete paths, and poor objectivity when facing complex and diverse chemical reactions, and there is an urgent need to improve its efficiency and reliability through automated methods.

[0003] Currently, there are some automated reaction path search methods, including Reaction Mechanism Generator (RMG), AFIR (Artificial Force Induced Reaction), NEB (Nudged Elastic Band), GSM (Growing String Method), etc. These methods have their own characteristics and are applicable to different types of chemical reactions. Although the existing automated reaction path search methods have made significant progress, there are still many deficiencies, such as dependence on initial guesses, high computational complexity, combinatorial explosion problems, lack of chemical inspiration, low prediction accuracy, insufficient reliability, poor interpretability of reaction paths, imperfect automatic error correction mechanisms, etc. Summary of the Invention

[0004] The purpose of the present invention is to overcome the deficiencies in the prior art, and provide a multi-scale heuristic automatic exploration system for chemical reaction networks and its system, to achieve efficient search, automatic verification and data analysis of reaction paths, and improve the reliability of results.

[0005] To achieve the above object, the present invention is implemented by the following technical solutions:

[0006] On the one hand, the present invention provides a multi-scale heuristic automatic exploration method for chemical reaction networks, comprising the following steps:

[0007] Step S1: Use a molecular structure processor to perform molecular conformation preprocessing on multiple input molecular structures participating in chemical reactions;

[0008] Step S2: Combine the decision rules inspired by chemical knowledge, and for the preprocessed molecular conformations, use multi-dimensional physical and chemical feature quantification analysis to automatically identify high-activity sites and pre-exclude inefficient paths, and use hierarchical screening to retain candidate reaction site pairs with polar matching and reasonable spatial proximity;

[0009] Step S3: After obtaining the candidate reaction site pairs, generate preliminary reaction path hypotheses by combining reaction rule templates or user-defined reaction directions, and use a "hierarchical screening - progressive optimization" calculation strategy to perform high-precision path optimization on the preliminary reaction path hypotheses to explore reaction paths;

[0010] Step S4: Perform path validity verification, real-time feedback optimization, and anomaly detection and recovery on the explored reaction paths, obtain reliable elementary reaction paths, and construct a complete reaction network and a convergence judgment mechanism based on a network expander;

[0011] Step S5: Store the final reaction network in a graph database, and at the same time generate a thermodynamic data table containing parameters such as activation free energy and pre-exponential factor, and perform visual output through an interactive analysis interface.

[0012] On the other hand, the molecular conformation preprocessing at least includes charge and spin state determination, geometric structure standardization, and initial conformation set generation, comprising the following steps:

[0013] Automatically assign charge states based on Hirshfeld charge distribution and determine the spin multiplicity through the parity check of the number of electrons and Hund's rule;

[0014] Bond length correction and meeting the allowable deviation and implicit addition of hydrogen atoms based on atomic hybridization type;

[0015] Generate initial conformations for each input molecule using the ETKDG algorithm of RDKit or force field optimization.

[0016] On the other hand, the step S1 further includes loading preset parameters, and the preset parameters include calculation accuracy level, path search depth, filter threshold, and scheduler to allocate computing resources.

[0017] On the other hand, in step S2, use an active site filter to perform multi-dimensional analysis on the preprocessed molecular conformations, including:

[0018] Extract electronic structure features based on local reactivity analysis, electrostatic complementarity analysis, and bond order analysis;

[0019] And perform steric effect evaluations including calculating buried volume, calculating Connolly surface accessibility, calculating spatial accessibility, and metal coordination space evaluation.

[0020] On the other hand, in step S2, use a polymer filter to exclude invalid clusters caused by insufficient intermolecular forces or incomplete structures, and the parameter settings are: the shortest intermolecular distance Exclude incomplete structures with more than 50% hydrogen bond breakage or bond angle deviation > 15°; exclude π-π stacking distances Of invalid stacking patterns.

[0021] On the other hand, in step S2, synchronously intervene and match the reaction rule library, and the reaction rule library includes element rules, distance rules, and template rules; the element rules first verify the number of d electrons of the metal center, the distance rules dynamically adjust the reaction atom spacing threshold, and match with the preset reaction templates in the template rule library.

[0022] On the other hand, in step S3, enter the coarse-grained screening stage, use the semi-empirical method GFN2-xTB for large-scale molecular dynamics pre-sampling, and screen out the 10 lowest-energy representative conformations through the hierarchical clustering algorithm;

[0023] Generate candidate paths based on the reaction rule system or user settings guidance, and use the semi-empirical method GFN2-xTB to quickly evaluate and pre-screen the candidate paths: screen out candidate paths with ΔE < 50 kcal / mol through constrained potential energy surface scanning, and use the GSM-SE algorithm to roughly locate the transition state, and then estimate the energy barrier and exclude paths with too high energy barriers;

[0024] Based on the high-precision quantum chemistry calculation framework, start the primary optimization L1 layer, the advanced optimization L2 layer, and the high-precision single-point energy and solvent effect correction L3 layer of the hierarchical optimization process respectively, and adaptively select calculation methods according to the system characteristics to integrate multiple transition state automatic search methods;

[0025] After obtaining the transition state structure, perform IRC calculations to verify the continuity of the reaction path.

[0026] On the other hand, the path validity verification of the explored reaction path includes structural verification, electronic state inspection, thermodynamic verification, and kinetic verification with multiple filters working together;

[0027] The structural verification uses bond length rationality inspection, C-C bond range: And virtual frequency detection is performed to ensure that the intermediate has no virtual frequencies and the transition state has only one virtual frequency; for the electronic state check, 〈S2〉 deviation < 0.1 is adopted to ensure no spin contamination in DFT calculations;

[0028] For the thermodynamic verification, microscopic reversibility test is adopted, with the ΔG difference between the forward and reverse reactions < 1 kcal / mol, and free energy continuity check, with ΔG between adjacent states < 30 kcal / mol; for the kinetic verification, the Eyring equation is used to verify the transition state, and the reaction rate formula is

[0029] If there is no experimental reference value, the reaction rate k is required to be < 4.0×10 -6 s -1 , that is, the half-life is 2 days;

[0030] If there is an experimental reference value, log(k pred / k ref ) < 0.5 is required.

[0031] On the other hand, based on the verified path, sub-reaction nodes are deduced, and energy pruning strategy and topological redundancy removal strategy are applied. When the double convergence condition is reached, the final reaction network is stored in the graph database, and at the same time, a thermodynamic data table containing parameters such as activation free energy and pre-exponential factor is generated.

[0032] On the other hand, the present invention provides a multi-scale heuristic chemical reaction network automatic exploration system, including the following modules:

[0033] Engine core module, which is used to overall coordinate and schedule the entire calculation process through the main controller. After the user submits a task through the interaction interface, the engine core first reads the parameter settings in the configuration manager, and then the scheduler allocates computing resources based on the task priority;

[0034] Calculation method module, configured with a semi-empirical method module and a quantum chemistry calculation module. The semi-empirical method module is used for rapid structural pre-optimization and preliminary reaction path scanning, and then the preliminary results are passed to the quantum chemistry calculation module for high-precision optimization;

[0035] Core function component module, configured with a molecular structure processor for processing molecular configurations and reaction mechanisms; a reaction processor for analyzing reaction mechanisms and verifying paths; a thermodynamics module for calculating energies and thermodynamic parameters; a kinetics module for calculating reaction rates and kinetic parameters;

[0036] Multiple filtering module, configured with an active site filter, a polymer filter, and a structure filter, for performing multiple filtering verifications after obtaining the elementary reaction paths at the quantum chemistry calculation level;

[0037] The reaction rule module, including element rules, distance rules, polarization rules, and template rules, is used to continuously supervise and guide the progress of calculations throughout the process;

[0038] The database module is used to write all calculation data into the configured structure database, reaction database, compound database, and calculation database through the result transmitter in response;

[0039] The reaction control module is configured with a network expander and a result transmitter. Based on the existing results, the network expander predicts the possibility of new reactions, and data is exchanged with other modules through the result transmitter.

[0040] Compared with the prior art, the beneficial effects achieved by the present invention: A multi-scale heuristic automatic exploration method and system for chemical reaction networks provided by the present invention. The core innovation lies in constructing a fully automatic chemical reaction network exploration system. By integrating chemical intelligence and computational intelligence, efficient mechanism analysis without manual intervention is achieved. Only the initial molecular structure participating in the reaction needs to be input by the user, and then the reaction path search, verification, and network construction can be automatically completed, and finally a visual and interpretable reaction mechanism map is output. The present invention utilizes a collaborative mechanism of chemical heuristic intelligent filtering and priority scheduling to significantly reduce the amount of invalid calculations through multi-level pre-screening and verification. Brief Description of the Drawings

[0041] Figure 1 It is a flowchart of traditional prior art means.

[0042] Figure 2 It is a block diagram of the multi-scale heuristic automatic exploration method and system for chemical reaction networks provided by the embodiments of the present invention.

[0043] Figure 3 It is a schematic diagram of a system interface provided by the embodiments of the present invention.

[0044] Figure 4 It is a schematic diagram of a molecular structure input module provided by the embodiments of the present invention.

[0045] Figure 5 It is a schematic diagram of a computing system configuration module provided by the embodiments of the present invention.

[0046] Figure 6 It is a schematic diagram of a calculation parameter configuration module provided by the embodiments of the present invention.

[0047] Figure 7 It is a schematic diagram of a filter system module provided by the embodiments of the present invention.

[0048] Figure 8 It is a schematic diagram of a reaction result analysis and visualization module provided by the embodiments of the present invention. Detailed Embodiments

[0049] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention and should not be used to limit the protection scope of the present invention.

[0050] Embodiment

[0051] As Figure 2 shown, in the embodiment of the present invention, a multi-scale heuristic automatic exploration method for chemical reaction networks is provided, which includes specific function implementation and data processing methods, and a modular and hierarchical design, and can efficiently and reliably complete the automatic exploration task of chemical reaction networks.

[0052] Step 1: The user submits the molecular structures of multiple participating chemical reactions through the interaction interface, and the system starts the molecular structure processor to perform standardized preprocessing on the input data. In this stage, the molecular charge state and spin multiplicity are automatically determined first, and then geometric structure standardization is performed, and the initial conformation set is generated using the ETKDG algorithm or force field optimization of RDKit.

[0053] After the preprocessing is completed, the configuration manager loads the preset parameters and user-defined parameters (including calculation accuracy level, path search depth, filter threshold), and the scheduler dynamically allocates CPU cores and memory according to the current hardware resources.

[0054] Step 1.1: Molecular standardization preprocessing

[0055] In this step, first, the input interface is configured. The user submits the molecular structure through the visualization interface, supporting formats such as SMILES and XYZ. Secondly, the molecular structure processor is executed, mainly including charge and spin state determination, geometric standardization, and initial conformation generation.

[0056] Specifically, in terms of charge and spin state determination, the charge state is automatically assigned based on the Hirshfeld charge distribution, and the spin multiplicity is determined through the parity check of the number of electrons and Hund's rule. In terms of geometric standardization, bond length correction (referring to the IUPAC standard database, allowing deviation) and implicit addition of hydrogen atoms (based on atomic hybridization type) are adopted. In terms of initial conformation generation, the initial conformation of each input molecule is generated using the ETKDG algorithm of RDKit or force field optimization.

[0057] Step 1.2: The configuration manager loads the preset parameters, and the specific parameter settings are as follows:

[0058] 1.2.1) Calculation accuracy level (default DFT functionals L1: B97-3c, L2: B3LYP, L3: wB97M-V).

[0059] 1.2.2) Path search depth (default 5-level expansion, maximum 10-layer reaction network).

[0060] 1.2.3) Filter threshold (bond length deviation automatically removed).

[0061] 1.2.4) The scheduler allocates computing resources (number of CPU cores and memory occupancy).

[0062] Step 2, use the heuristic filtering module to intelligently identify reaction sites.

[0063] After completing the molecular conformation preprocessing, the heuristic filtering module is started. Through the quantitative analysis of multi-dimensional physicochemical characteristics and combined with the decision rules inspired by chemical knowledge, the automatic identification of high-activity sites and the pre-exclusion of inefficient paths are realized. Through a hierarchical screening system, only candidate reaction site pairs with polar matching and reasonable spatial proximity are retained, so as to drive the candidate site pairs closer in the subsequent path exploration module to guide the occurrence of chemical reactions.

[0064] The active site filter is started to quantitatively analyze the electronic structure characteristics and the spatial accessibility of atoms / sites, mark the sites that may participate in the reaction, and exclude the regions that cannot participate in the reaction due to excessive steric hindrance. The aggregate filter works synchronously to exclude thermodynamically unstable or kinetically inert aggregate forms.

[0065] The multiple filtering verification mechanism uses chemically inspired intelligent driving, step-by-step screening strategies, multi-dimensional rationality checks, and automated quality control to effectively identify candidate reaction site pairs.

[0066] At the same time, the reaction rule system intervenes. Specifically, the element rule first verifies the number of d electrons of the metal center, the distance rule dynamically adjusts the reaction atom spacing threshold, and the preset reaction templates in the template rule library (covering 32 types of mechanisms such as C-H activation and coupling reactions) are matched. Secondly, in this embodiment, the reaction rule system integrates physical and chemical knowledge with databases, multi-level rule systems, and flexible rule extension interfaces.

[0067] Generate the atomic pair priority list P ij , and screen out high-priority atomic pairs (P ij > 0.5) to allocate 80% of the computing resources (CPU cores + high-speed memory), medium-priority tasks (0.3 <P ij ≤ 0.5) to allocate 15% of the resources, and low-priority (P ij ≤ 0.3) tasks to allocate 5% of the resources.

[0068] The specific process steps and technical parameters are as follows.

[0069] Step 2.1: The active site filter conducts multi-dimensional analysis on the preprocessed conformation, mainly including extraction of electronic structure features and evaluation of steric effects.

[0070] Specifically, in the extraction of electronic structure features, local reactivity analysis, electrostatic complementarity analysis, and bond order analysis are respectively carried out. Among them, in the local reactivity analysis, the Fukui function (f + =ρ N+1 -ρ N ) and (f - =ρ N -ρ N-1 ) are calculated. Atoms with f + >0.2 are marked as nucleophilic sites (such as electron-rich olefin carbon atoms), and atoms with f - >0.15 are marked as electrophilic sites (such as electron-deficient carbonyl oxygen atoms). The electrostatic complementarity analysis is based on the CHELPG charge partitioning method to calculate the electrostatic potential (ESP) distribution on the molecular surface. regions are marked as electron-rich regions (such as aromatic ring π electron clouds), regions are marked as electron-deficient regions (such as high-valent metal centers). In the bond order analysis, sites with a Wiberg bond order > 0.3 are judged as potential active sites.

[0071] In the evaluation of steric effects, it mainly includes calculation of the buried volume, calculation of Connolly surface accessibility, calculation of steric accessibility, and evaluation of the metal coordination space.

[0072] Among them, when calculating the buried volume, with the target atom as the center, the van der Waals occupancy of other atoms within the spherical shell around it is calculated to characterize the spatial exposure degree of the site. The formula for calculating the buried volume V bunied of atom A using the Sambvca2.0 algorithm is:

[0073]

[0074] If the buried volume V bunied > 70% of the site, it is marked as an inaccessible site.

[0075] Calculate the Connolly surface accessibility (probe radius ), and exclude sites with .

[0076] Calculate the steric accessibility and exclude atoms with a Voronoi volume of .

[0077] In the evaluation of the metal coordination space, for transition metal complexes (such as Pd catalysts), calculate the Tolman cone angle:

[0078]

[0079] Among them, r ligand is the van der Waals radius of the ligand, and R M-L is the metal-ligand bond length.

[0080] Exclude the metal center with the cone angle θ T > 180° (steric crowding prevents coordination).

[0081] Step 2.2: The polymer filter excludes invalid clusters caused by insufficient intermolecular forces or incomplete structures.

[0082] Among them, the invalid cluster exclusion parameter is set as: the shortest intermolecular distance (exceeding the typical non-bonding interaction range). The hydrogen bond network detection parameter is set as: exclude incomplete structures with more than 50% hydrogen bond breakage or bond angle deviation > 15°. The parameter for organic molecules or fragments is set as: exclude the invalid stacking mode with the π-π stacking distance .

[0083] Step 2.3: The reaction rule matching system intervenes synchronously.

[0084] Among them, the reaction rule library matching includes element rule verification, distance rule screening, and template rule matching. Specifically, the element rule verification needs to check the number of d electrons of the metal center (such as d 8 configuration triggers the oxidative addition rule) and verify the coordination saturation (coordination number < 18 electron rule). The parameter setting for distance rule screening is to set the dynamic reaction radius (T is the temperature) and exclude candidate site pairs with atomic spacing exceeding (R). The template rule matching presets reaction modes (covering 32 types of organic / metal-organic reaction templates). For example, to trigger the oxidative addition rule condition, the oxidation state of the metal atom needs to increase by +2.

[0085] Table 1 Reaction Rules

[0086]

[0087] Step 3, Multi-scale calculation engine and path exploration.

[0088] After obtaining the list of candidate reaction atom pairs through the multiple filtering system, the calculation engine and path exploration system generate preliminary reaction path hypotheses by combining reaction rule templates or user-defined reaction directions, and adopt a "hierarchical screening - progressive optimization" calculation strategy for reaction path exploration. The multi-level calculation strategy takes into account both calculation efficiency and accuracy, combines semi-empirical and high-precision quantum chemistry methods, and uses an intelligent method selection mechanism to adaptively adjust the calculation accuracy.

[0089] First, enter the coarse-grained screening stage. The semi-empirical method module (GFN2-xTB) first performs large-scale molecular dynamics pre-sampling and screens out the 10 representative conformations with the lowest energy through the hierarchical clustering algorithm. The semi-empirical method module then quickly evaluates these paths: screens out candidate paths with ΔE < 50 kcal / mol through constrained potential energy surface scanning, roughly locates the transition state using the GSM-SE algorithm, and then estimates the energy barrier and excludes paths with too high energy barriers.

[0090] Then, the system switches to the high-precision quantum chemistry calculation framework. The quantum chemistry calculation module starts the hierarchical optimization process (L1, L2, L3), adaptively selects the calculation method according to the system characteristics, and integrates multiple methods for automatic search of transition states. During this process, the scheduler monitors the calculation process in real time and ensures the path continuity through IRC verification and algorithm switching.

[0091] The specific process and technical parameters are as follows:

[0092] Step 3.1: Molecular conformation sampling and coarse screening. Use the semi-empirical method module (GFN2-xTB) to perform molecular dynamics pre-sampling and conformational clustering analysis respectively.

[0093] Among them, the parameter settings for molecular dynamics pre-sampling are: running 50 ps NVT simulation (300 K, step size 1 fs) and generating a trajectory file containing 200 conformations. The technical parameter settings for conformational clustering analysis are hierarchical clustering based on the RMSD matrix (cutoff value ) and retaining the 10 representative conformations with the lowest energy.

[0094] Step 3.2: Quick path pre-screening, which is executed separately according to the generation of preliminary screening paths and the semi-empirical method module.

[0095] Among them, the generation of preliminary screening paths uses the reaction rule system or user-defined guidance to generate candidate paths. The technical parameter settings are: applying template rule matching (such as oxidative addition / reductive elimination), polarization rule to predict the direction of electron transfer (judged as valid when μ > 0.5 D), and distance rule to verify atomic proximity (the distance between reaction atoms ).

[0096] The execution of the semi-empirical method module mainly includes the following steps:

[0097] 3.2.1) Adopt constrained potential energy surface scanning (path pre-scanning). Specifically, scan along the preset reaction coordinate (bond stretching / angle bending) with a bond length or an angle of 5° as the step size and screen out candidate paths with ΔE < 50 kcal / mol.

[0098] 3.2.2) Use the GSM-SE (String Method Semi-Empirical) algorithm and set the convergence condition: maximum atomic force

[0099] 3.2.3) Perform energy pre-estimation to calculate the path pruning strategy. Configure the technical parameters as follows:

[0100] Energy barrier threshold: (ΔE≠=E TS,xTB -E R,xTB ), if ΔE ≠ >50 kcal / mol, terminate this path; exceptions for exothermic reactions, if ΔE<-5 kcal / mol, the threshold is increased to 70 kcal / mol.

[0101] Step 3.3: High-precision path optimization.

[0102] The quantum chemistry calculation module starts hierarchical calculations, respectively starts the primary optimization L1 layer, the advanced optimization L2 layer, and the high-precision single-point energy and solvent effect correction L3 layer of the hierarchical optimization process, and adaptively selects calculation methods according to the system characteristics to integrate multiple transition state automatic search methods. After obtaining the transition state structure, perform IRC calculations to verify the continuity of the reaction path.

[0103] Among them, the functional of the primary optimization (L1 layer): B97-3c, basis set: mTZVP.

[0104] In the advanced optimization (L2 layer), the system adaptively selects a suitable calculation method according to the characteristics of the input molecular system:

[0105] Pure organic system: B3LYP-D3 functional / def2-TZVP basis set.

[0106] System containing a single transition metal: PBE0-D3 functional / def2-SVP / def2-TZVP mixed basis set.

[0107] System containing multiple metals: PBE-D3 functional / def2-SVP / def2-TZVP mixed basis set.

[0108] In the high-precision single-point energy and solvent effect correction (L3 layer), use the SMD implicit solvent model to consider the solvent effect, and select different high-precision methods for single-point energy calculation calibration according to the system size and characteristics. The specific parameter settings are: number of atoms > 150: wB97M-V / def2-TZVP; 80 < number of atoms < 150: PWPB95-D4 / def2-TZVP; number of atoms < 50: DLPNO-CCSD(T) / def2-TZVPP with normalPNO.

[0109] Integrate three algorithms for automatic search of transition states in L1 and L2 layers:

[0110] (a) NEB algorithm (default): 20 image points, spring constant

[0111] (b) GSM-DE algorithm: string dynamic increase and decrease mechanism;

[0112] (c) Dimer method: displacement step size Curvature convergence threshold

[0113] After obtaining the transition state structure, perform IRC calculation to verify the continuity of the reaction path.

[0114] Step 4, Dynamic verification and feedback

[0115] After obtaining the elementary reaction path at the quantum chemistry calculation level, the multiple filtering system conducts strict verification, which can achieve adaptive closed-loop evolution, real-time feedback of calculation results, automatic parameter optimization, and intelligent path selection, etc.

[0116] The structure filter first checks the rationality of bond lengths and ensures that the molecular structure has no extra imaginary frequencies. The electronic state checks the spin quantum number to ensure that there is no spin contamination in DFT calculations.

[0117] The thermodynamics module calculates the free energy change ΔG, verifies the free energy continuity (ΔG < 30 kcal / mol for adjacent states) and microscopic reversibility (the difference in ΔG between forward and reverse reactions < 1 kcal / mol). The kinetics module checks the consistency of the reaction rate with the experimental value (if there is an experimental reference value) through the Eyring equation.

[0118] The paths that fail to pass the verification will trigger an automatic error correction mechanism: the paths with abnormal structures are sent back to the molecular structure processor for re-optimization, and the paths with abnormal energies are upgraded to the high-precision layer (L3) for re-calculation.

[0119] Specifically, it includes the following processes and technical parameters.

[0120] Step 4.1: Path validity verification.

[0121] During the path validity verification process, a collaborative working mechanism of a multiple filtering system for structure verification, electronic state inspection, thermodynamics verification, and kinetics verification is adopted.

[0122] Specifically, the structure verification adopts and sets the rationality check of bond lengths (such as the C-C bond range: ), and imaginary frequency detection (ensuring that there is no imaginary frequency for the intermediate and only one imaginary frequency for the transition state). In the electronic state check, set the deviation of 〈S2〉 < 0.1 (ensuring no spin contamination in DFT calculations). Thermodynamic verification includes microscopic reversibility test and free energy continuity check. The parameter for the microscopic reversibility test is that the ΔG difference between the forward and reverse reactions < 1 kcal / mol, and the parameter for the free energy continuity check is that the ΔG between adjacent states < 30 kcal / mol. The kinetic verification uses the Eyring equation to verify that the transition state parameter value is: reaction rate formula: Among them, if there is no experimental reference value, it is required that the reaction rate k < 4.0×10 -6 s -1 , that is, the half-life is 2 days. If there is an experimental reference value, it is required that log(k pred / k ref ) < 0.5.

[0123] Step 4.2: Real-time feedback optimization.

[0124] In the optimization, a scheduler dynamic adjustment strategy is adopted, which specifically includes the following:

[0125] 4.2.1) When the IRC path is detected to be broken, automatically switch to the Dimer method for re-search;

[0126] 4.2.2) When the energy difference is greater than the threshold, trigger the high-precision layer (L3) calculation;

[0127] 4.2.3) Dynamic weight adjustment mechanism: When the calculation of the preferred path (high priority) fails (such as the TS does not converge), reduce the weight of the original task (Δ weight = -0.2 / time), increase the weight of the sub-optimal path (Δ weight = +0.3), and optimize the parameters based on the failure reason (such as increasing the NEB spring constant by 5%).

[0128] Step 4.3 Anomaly detection and recovery.

[0129] When SCF does not converge, automatically switch to algorithms such as DIIS / Pulay; and when there are too many structural imaginary frequencies, trigger methods such as frequency projection correction to eliminate the excess imaginary frequencies.

[0130] Step 5, Network expansion and convergence.

[0131] After obtaining a reliable elementary reaction path, the network expander starts to construct a complete reaction network. The system derives sub-reaction nodes based on the verified path (for example, triggering the β-H elimination rule after C-H activation), and applies an energy pruning strategy (retaining paths with ΔG < 35 kcal / mol) and a topological redundancy elimination strategy (excluding similar paths). When the double convergence condition (no lower energy path is found in 3 consecutive iterations and the number of newly added nodes < 5% of the total number of nodes) is reached, the system stores the final reaction network in the graph database and generates a thermodynamic data table containing parameters such as the activation free energy and the pre-exponential factor.

[0132] Step 5.1, Reaction network construction.

[0133] Use the network expander to start constructing the complete reaction network. Among them, the network expander performs the following steps:

[0134] 5.1.1) Path branch generation, deduce possible sub-reactions through the reaction rule system, such as triggering the β-H elimination rule after C-H activation.

[0135] 5.1.2) Energy pruning strategy, retain paths with ΔG < 35 kcal / mol and exclude redundant branches with existing paths of.

[0136] Step 5.2: Convergence judgment mechanism.

[0137] Set energy convergence and topological convergence in the double convergence criterion. The specific parameters are: set no lower energy path is found in 3 consecutive iterations in energy convergence, set the number of newly added nodes < 5% of the total number of nodes in topological convergence, and terminate the calculation when any of the above conditions is met.

[0138] When the termination condition is triggered, output the final reaction network diagram and generate a thermodynamic parameter table ( activation energy, etc.).

[0139] Step 6, Data management and output. The finally calculated results will be stored in the database and visualized and output through an interactive analysis interface.

[0140] Step 6.1: Structured storage. All calculation data is written into the database system through the result transmitter.

[0141] First, construct the database system design. The structure database stores the optimized configuration in XYZ format (precision ). The calculation database archives energy data and kinetic parameters for hierarchical storage including calculation method dependencies, and the reaction database records path topological relationships to store into the reaction network diagram.

[0142] Step 6.2: Visualization output.

[0143] An interactive analysis interface is used to provide the following information and support export to local: reaction network diagram (displayed based on the networkx library), free energy potential energy surface diagram (displayed based on the matplotlib library), three-dimensional molecular structure (displayed based on the py3Dmol library), and visualization of electronic structure and steric hindrance (displayed based on calculations using the Multiwfn software).

[0144] As Figure 2 shown, the embodiment of the present invention also provides a multi-scale heuristic automatic exploration chemical reaction network system, which is designed with a modular hierarchical architecture and mainly includes the following core modules. Each module can run independently or collaboratively to achieve the full-process calculation from molecular input to reaction network construction and result output.

[0145] Specifically, it includes the following modules (or systems):

[0146] Engine core module, which is used to overall coordinate and schedule the entire calculation process through the main controller. After the user submits a task through the interactive interface, the engine core first reads the parameter settings (such as calculation accuracy level, filter threshold, etc.) in the configuration manager, and then the scheduler allocates computing resources based on the task priority.

[0147] Among them, the main controller is responsible for overall task scheduling and resource allocation, managing data interaction between modules, and controlling the execution of the calculation process. The configuration manager is responsible for managing system default parameters, processing user-defined settings, and optimizing runtime configurations. The user interaction interface provides a task submission interface, displays the calculation progress and results, and supports real-time parameter adjustment.

[0148] Calculation method module, which is configured with a semi-empirical method module and a quantum chemistry calculation module. The semi-empirical method module is used for rapid structural pre-optimization and preliminary reaction path scanning, and then the preliminary results are passed to the quantum chemistry calculation module for high-precision optimization.

[0149] Among them, the semi-empirical method module uses GFN-xTB for rapid structural optimization, and can provide rapid vibration analysis, molecular dynamics conformational search, and reaction path scanning (coordinate-driven / GSM-SE). The quantum chemistry calculation module uses DFT / double hybrid method for structural optimization, and provides vibration frequency analysis, transition state search (NEB / Dimer / GSM-DE), IRC calculation, and high-precision single-point energy calculation. The seamless connection between the semi-empirical method module and the quantum chemistry calculation module ensures the continuity from the initial conformation to the final transition state and reduces the amount of invalid calculations.

[0150] The preliminary optimization results of the semi-empirical method module are transmitted to the compound processor for structure analysis. The high-precision results of the quantum chemistry calculation module are used for the mechanism verification of the reaction processor. The thermodynamic and kinetic parameters of the verified reaction path are used by the kinetic module and the thermodynamic module to calculate the reaction rate and energy. Reaction rules are applied in path scanning and transition state search to ensure the chemical rationality of the path. The transition state structures generated by high-precision optimization are written into the database through the result transmitter, triggering the network expander to deduce derivative paths.

[0151] The core functional component module is configured with a molecular structure processor for processing molecular configurations and reaction mechanisms, a reaction processor for analyzing reaction mechanisms and verifying paths, a thermodynamic module for calculating energy and thermodynamic parameters, and a kinetic module for calculating reaction rates and kinetic parameters.

[0152] Among them, the scheduler is responsible for task priority management, computing resource allocation, and task status monitoring. The molecular structure processor is responsible for structure standardization, conformational analysis, and geometric parameter optimization. The reaction processor provides reaction mechanism analysis, path verification, and energetics calculation. The thermodynamic module provides energy calculation, thermodynamic parameter analysis, and equilibrium constant calculation. The kinetic module provides reaction rate calculation, activation energy analysis, and kinetic parameter estimation.

[0153] The reaction processor combines the thermodynamic and kinetic modules to ensure that the generated path conforms to both thermodynamic principles and kinetic feasibility, forming a complete closed-loop for reaction path verification. The structure analysis results of the compound processor are used for cluster analysis of the polymer filter, and the mechanism verification results of the reaction processor are used for filtering active sites and elementary steps. The thermodynamic parameters and kinetic data are written into the reaction database through the result transmitter to support the energy pruning strategy of the subsequent network expander.

[0154] The multiple filtering module is configured with an active site filter, a polymer filter, and a structure filter for performing multiple filtering validations after obtaining the elementary reaction path at the quantum chemistry calculation level.

[0155] Among them, the polymer filter is responsible for cluster structure analysis, intermolecular interaction evaluation, and stability prediction. The active site filter is responsible for reaction site identification, electronic structure analysis, and bonding property evaluation. The elementary step filter provides reaction step verification, energetics screening, and mechanism rationality check. The structure filter provides conformational rationality check, bond length and bond angle verification, and steric hindrance analysis. The active site filter and the polymer filter work together to ensure that the selected candidate site pairs have polar matching and reasonable spatial proximity. The elementary step filter and the structure filter further verify the rationality of the path, reduce invalid calculations, and improve the calculation efficiency and result reliability.

[0156] Reaction rules are applied in active site screening to ensure that the screening results are consistent with chemical knowledge; paths that fail filtering are fed back to the main controller for task rescheduling.

[0157] Reaction rule modules, including element rules, distance rules, polarization rules, and template rules, are used to continuously supervise and guide the calculation throughout the process. The various rule modules work together to continuously supervise and guide the calculation process to ensure that the generated path is both in line with chemical principles and practical.

[0158] Among them, element rules are used for atomic valence state inspection, electronic configuration verification and bonding rule verification. Distance rules are used for atomic spacing inspection, spatial configuration verification and molecular packing analysis. Polarization rules are used for charge distribution analysis, polarization effect evaluation and electron transfer prediction. Template rules are used for reaction type matching, reaction pattern recognition and mechanism template application.

[0159] The reaction rule module works in conjunction with the multiple filtering module to ensure that the selected paths are reasonable at both the chemical and physical levels, improving the accuracy of path prediction. The reaction rule module works in conjunction with the calculation method module to apply rules in path scanning and transition state search to guide the calculation process. The reaction rule module interacts with the reaction control system, and the rule matching results are used for path derivation and optimization of the network expander.

[0160] The database module, as the core of data storage and management, ensures the complete record and efficient access of all calculation results. It is used to write all calculation data into the corresponding configured structure database, reaction database, compound database and calculation database through the result transmitter. Combined with the result transmitter, it realizes data exchange with other modules and supports dynamic expansion and update of reaction networks.

[0161] Integrating with the reaction control system, the network extender reads the historical path from the database to guide the derivation and optimization of the new path. Interacting with the result transmitter, the asynchronous exchange and storage of data is realized through the result transmitter.

[0162] The reaction control module is equipped with a network expander and a result transmitter. It predicts new reaction possibilities based on existing results through the network expander and exchanges data with other modules through the result transmitter. Through close integration with the database module, the reaction control module realizes dynamic expansion of reaction paths and real-time update of data, supporting continuous optimization of the system.

[0163] In collaboration with the core functional components, the path optimization results of the network extender are fed back to the reaction processor for mechanism verification. Interacting with the database system, the optimization results are stored in the corresponding database through the result transmitter to support subsequent analysis and application.

[0164] Through calculation verification, this system demonstrates significant advantages in multiple actual cases. It can quickly and accurately predict and analyze complex chemical reaction networks. Through efficient automation, it reduces manual intervention and improves calculation efficiency; through intelligent screening, it reduces ineffective calculations and lowers calculation costs; through cross-scale methods, it ensures the accuracy and reliability of calculation results and the rationality of paths. The present invention provides powerful computational tool support for research related to chemical reaction mechanisms and predictions, etc.

[0165] Table 2

[0166]

[0167] Specifically, the improvement in calculation efficiency is manifested as a 60 - 150% increase in calculation speed, a 40% increase in resource utilization rate, and a 60% increase in parallel calculation efficiency. The improvement in automation level is manifested as a 90% reduction in manual intervention, a 200% increase in task processing ability, and a 40% increase in reaction prediction accuracy. The improvement in result reliability is manifested as an 80% reduction in error rate, a 70% increase in result reproducibility, and a 50% increase in prediction accuracy. The expansion of the application scope is manifested as supporting more reaction types, being applicable to more complex systems, and handling larger reaction networks. The improvement in user experience is manifested as improved operation simplicity, enhanced result visualization, and faster interaction response.

[0168] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art of this technology, without departing from the technical principle of the present invention, several improvements and modifications can still be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A multi-scale heuristic chemical reaction network automatic exploration method, characterized in that: The following steps are involved: Step S1: using a molecular structure processor to perform molecular conformation preprocessing on a plurality of input molecular structures involved in a chemical reaction; Step S2: Combining the judgment rules inspired by chemical knowledge, the pre-processed molecular conformations are subjected to multi-dimensional physicochemical feature quantitative analysis to automatically identify highly active sites and pre-exclude inefficient pathways, and hierarchical screening is used to retain candidate reaction site pairs with polarity matching and reasonable spatial proximity; Step S3: After obtaining the candidate reaction site pairs, a preliminary reaction path hypothesis is generated in combination with the reaction rule template or the user-defined reaction direction, and the preliminary reaction path hypothesis is optimized with high precision using the calculation strategy of "hierarchical screening-progressive optimization" to explore the reaction path; Step S4: perform path validity verification, real-time feedback optimization, and anomaly detection and recovery on the explored reaction path, obtain a reliable elementary reaction path, and build a complete reaction network and convergence judgment mechanism based on the network expander; Step S5: The final reaction network is stored in a graph database, and a thermodynamic data table containing parameters such as activation free energy and pre-exponential factor is generated, and visualized output is performed through an interactive analysis interface.

2. The multi-scale heuristic chemical reaction network automatic exploration method according to claim 1, characterized in that: The molecular conformation preprocessing at least includes charge and spin state determination, geometric structure standardization and initial conformation set generation, including the following steps: Automatically assign charge states based on Hirshfeld charge distribution and determine spin multiplicity through electron parity check and Hund's rule; Bond length correction and meet the allowable deviations and implicit addition of hydrogen atoms based on atomic hybridization type; Initial conformations were generated for each input molecule using the ETKDG algorithm of RDKit or force field optimization.

3. The multi-scale heuristic chemical reaction network automatic exploration method according to claim 2, characterized in that: The step S1 also includes loading preset parameters, wherein the preset parameters include calculation accuracy level, path search depth, filter threshold, and scheduler allocation of computing resources.

4. The multi-scale heuristic chemical reaction network automatic exploration method according to claim 3, characterized in that: In step S2, the pre-processed molecular conformation is subjected to multi-dimensional analysis using active site filters, including: Extract electronic structure features based on local reactivity analysis, electrostatic complementarity analysis, and bond-order analysis; As well as steric effect evaluation including calculation of buried volume, calculation of Connolly surface accessibility, calculation of steric accessibility and evaluation of metal coordination space.

5. The multi-scale heuristic chemical reaction network automatic exploration method according to claim 4, characterized in that: In step S2, the polymer filter is used to exclude invalid clusters caused by insufficient intermolecular forces or incomplete structures. The parameters are set as: Incomplete structures with more than 50% hydrogen bond breakage or bond angle deviation >15° were excluded; Invalid stacking mode.

6. The multi-scale heuristic chemical reaction network automatic exploration method according to claim 5, characterized in that: In step S2, a reaction rule library is synchronously intervened and matched, and the reaction rule library includes element rules, distance rules and template rules; the element rule first verifies the number of d electrons in the metal center, the distance rule dynamically adjusts the threshold of the reaction atomic spacing, and the preset reaction template in the template rule library is matched.

7. The multi-scale heuristic chemical reaction network automatic exploration method according to claim 6, characterized in that: In step S3, the coarse-grained screening stage is entered, and large-scale molecular dynamics pre-sampling is performed using the semi-empirical method GFN2-xTB, and the 10 representative conformations with the lowest energy are screened out using the hierarchical clustering algorithm; Based on the reaction rule system or user-defined guidance, candidate pathways are generated, and the semi-empirical method GFN2-xTB is used to quickly evaluate and pre-screen the candidate pathways: candidate pathways with ΔE < 50 kcal / mol are screened through constrained potential energy surface scanning, and the GSM-SE algorithm is used to roughly locate the transition state, and then the energy barrier is estimated and pathways with too high energy barriers are excluded; Based on the high-precision quantum chemical calculation framework, the primary optimization L1 layer, advanced optimization L2 layer and high-precision single-point energy and solvent effect correction L3 layer of the hierarchical optimization process are started respectively, and the calculation method is adaptively selected according to the system characteristics to integrate a variety of transition state automatic search methods; After obtaining the transition state structure, IRC calculations were performed to verify the continuity of the reaction path.

8. The multi-scale heuristic chemical reaction network automatic exploration method according to claim 7, characterized in that: Verification of the validity of the explored reaction pathway includes structural verification, electronic state inspection, thermodynamic verification and kinetic verification through the coordinated work of multiple filters; The structure verification uses a bond length rationality check, CC bond range: And imaginary frequency detection to ensure that the intermediate has no imaginary frequency and the transition state has only one imaginary frequency; the electronic state check uses 〈S2〉 deviation <0.1 to ensure that the DFT calculation has no spin contamination; The thermodynamic verification uses a microscopic reversibility test, the difference between the forward and reverse reactions ΔG is <1 kcal / mol, and the free energy continuity check, the adjacent state ΔG is <30 kcal / mol; the kinetic verification uses the Erying equation to verify the transition state, and the reaction rate formula is If there is no experimental reference value, the reaction rate k is required to be less than 4.0×10 -6 s -1 , i.e. half-life is 2 days; If there is an experimental reference value, log(k pred / k ref )<0.

5.

9. The multi-scale heuristic chemical reaction network automatic exploration method according to claim 8, characterized in that: Sub-reaction nodes are derived based on the verified paths, and energy pruning strategy and topological redundancy removal strategy are applied. When the dual convergence conditions are met, the final reaction network is stored in the graph database, and a thermodynamic data table containing parameters such as activation free energy and pre-exponential factor is generated.

10. A multi-scale heuristic chemical reaction network automatic exploration system, characterized in that: Includes the following modules: The engine core module is used to coordinate and schedule the entire computing process through the main controller. When the user submits a task through the interactive interface, the engine core first reads the parameter settings in the configuration manager, and then the scheduler allocates computing resources based on the task priority; The calculation method module is equipped with a semi-empirical method module and a quantum chemical calculation module. The semi-empirical method module is used to perform rapid structural pre-optimization and preliminary reaction path scanning, and then the preliminary results are passed to the quantum chemical calculation module for high-precision optimization; The core functional component module is equipped with a molecular structure processor for processing molecular configurations and reaction mechanisms; a reaction processor for analyzing reaction mechanisms and verifying pathways; a thermodynamic module for calculating energy and thermodynamic parameters; and a kinetic module for calculating reaction rates and kinetic parameters. A multiple filtering module, configured with active site filters, polymer filters, and structure filters, is used to perform multiple filtering verification after obtaining the elementary reaction paths at the quantum chemical calculation level; Reaction rule modules, including element rules, distance rules, polarization rules, and template rules, are used to continuously supervise and guide the calculations throughout the process; A database module, used for writing all calculation data into the corresponding configured structure database, reaction database, compound database and calculation database through the result transmitter; The reaction control module is equipped with a network expander and a result transmitter. The network expander predicts the possibility of new reactions based on existing results, and the result transmitter exchanges data with other modules.