Method and system for analyzing enzymatic synthesis mechanism of N-nitrosourea compound

By constructing an AI-assisted multi-scale computing framework and combining MD simulation and QM/MM calculation, the problem of the difficulty in resolving the enzymatic synthesis mechanism of N-nitrosourea group in existing technologies has been solved, and atomic-level analysis and quantitative description of the enzymatic synthesis mechanism have been achieved.

CN122067631AActive Publication Date: 2026-05-19SHANDONG UNIV OF TRADITIONAL CHINESE MEDICINE
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANDONG UNIV OF TRADITIONAL CHINESE MEDICINE
Filing Date
2026-04-20
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing technologies are insufficient to accurately describe the enzymatic synthesis mechanism of N-nitrosourea groups. Crystallographic and kinetic experimental analysis methods are insufficient to capture the evolution of transient intermediates and transition states. Kinetic simulation methods are insufficient to handle electronic structure and protein environmental effects, making it difficult to evaluate the enzymatic synthesis mechanism at multiple scales.

Method used

A multi-scale computational framework was adopted, which includes AI-assisted quantitative determination of protonation state, automatic conformation screening based on MD simulation, generative AI prediction of reaction pathways, and QM/MM full pathway verification. Combined with GNN model, the enzyme-catalyzed synthesis mechanism was quantitatively analyzed. By predicting substrate pKa value, constructing complex model, molecular dynamics simulation and quantum mechanical calculation, the optimal reaction pathway was screened and the enzyme-catalyzed synthesis mechanism of N-nitrosourea group was quantitatively analyzed.

Benefits of technology

Atomic-level analysis of the enzymatic synthesis mechanism of N-nitrosourea groups was achieved, providing a complete quantitative relationship from the microenvironment of the enzyme active site to the reaction activity, and revealing the complete mechanism of enzymatic synthesis of N-nitrosourea pharmacophores.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122067631A_ABST
    Figure CN122067631A_ABST
Patent Text Reader

Abstract

The invention discloses a method and system for analyzing an enzymatic synthesis mechanism of N-nitrosourea compounds, and relates to the technical field of artificial intelligence, computational chemistry and synthetic biology. The method comprises the steps that firstly, the pKa value of each ionizable site of a substrate is predicted, and structural optimization and free energy calculation are conducted in combination with a quantum chemistry method; the protonation state of the substrate is quantitatively determined in the process; then molecular dynamics simulation is carried out to define an optimal reaction conformation and serve as an initial conformation, a plurality of potential reaction paths are obtained based on the initial conformation, QM / MM calculation is combined, and the hydrogen migration-attack path under the optimal reaction conformation is defined from the atomic level in the process, so that an enzyme activity center microenvironment characteristic to a substance protonation state is established, and a hydrogen migration-attack path under the optimal reaction conformation is obtained. The complete quantitative relation from the reaction activity to the reaction activity is realized, and the complete mechanism of N-nitrosourea pharmacophore enzymatic synthesis is disclosed from the atomic level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence, computational chemistry and synthetic biology, and in particular to a method, system, device and medium for elucidating the enzymatic synthesis mechanism of N-nitrosourea compounds. Background Technology

[0002] Streptozotocin (SZN) has held a significant position in medicine and life sciences since its discovery in the 1950s. It is currently the only chemotherapy drug approved by the U.S. Food and Drug Administration (FDA) for metastatic pancreatic islet cell carcinoma and also has irreplaceable value in the construction of animal models of type 1 diabetes. The N-nitrosourea group in its molecule is a key structural domain for exerting its biological activity. It can induce apoptosis or necrosis by alkylating DNA bases and activating cellular DNA damage response pathways, making it a classic probe in the core area of ​​tumor biology research.

[0003] Although the biological functions and clinical value of streptozotocin are widely recognized, the biosynthetic mechanism of its N-nitrosourea group has long been controversial. Early hypothesized non-enzymatic pathways (such as the reaction of nitrite with urea fragments under acidic conditions) lacked experimental evidence. In recent years, Balskus's group, through isotope labeling experiments and enzymatic analysis, confirmed that the formation of the N-nitrosourea group is entirely dependent on an enzymatic pathway, with the N-nitroso nitrogen atom derived from L-arginine, not exogenous nitrite. They also identified the key enzyme system SznE-FHJK, in which SznF plays a central catalytic role in the reaction. Its central domain is responsible for mediating L-NMA (… N ω N-methyl-L-arginine) δ Bit and N ω A two-step consecutive hydroxylation reaction occurs at the position to generate L-DHMA ( N δ , N ω -dihydroxy- N ω -Methyl-L-arginine); while the cupin domain promotes the N atom migration and oxidative rearrangement of L-DHMA to generate L-NHMA (-methyl-L-arginine). N δ -hydroxy- N ω -methyl- N ω -nitroso-L-citrulline).

[0004] However, due to the difficulty in quantitatively describing the enzymatic synthesis mechanism of N-nitrosourea groups, current research typically employs crystallographic and kinetic experimental analysis, as well as kinetic simulation. However, crystallographic and kinetic experimental analysis struggles to capture the evolution of transient intermediates and transition states, while kinetic simulation cannot simultaneously address electronic structure and protein environmental effects, making it difficult to accurately describe the reaction pathway. Therefore, current methods are insufficient to accurately describe the evolution pathways and processes of intermediate and transition state reactions, hindering multi-scale evaluation of the enzymatic synthesis mechanism of N-nitrosourea groups. Summary of the Invention

[0005] This invention provides a method and system for elucidating the enzymatic synthesis mechanism of N-nitrosourea compounds, which can solve the problems existing in the prior art.

[0006] This invention provides a method for elucidating the enzymatic synthesis mechanism of N-nitrosourea compounds, comprising the following steps: Obtain structural data of the superfamily enzymes and substrates used for the catalytic synthesis of N-nitrosourea groups; based on the molecular structural features in the substrate structural data and the microenvironment characteristics of the enzyme active site of the superfamily enzymes, predict the p-values ​​of each ionizable site of the substrate. K a Value; based on p of each ionizable site of the substrate K a The protonation state of the substrate was determined by using quantum chemical methods to optimize the structure and calculate the free energy in gas phase and implicit solvent environments, respectively. We obtained the protein crystal structure data of the superfamily enzymes and constructed a complex model including the superfamily enzymes and substrates based on the protein crystal structure data of the superfamily enzymes and substrates. Based on the protonation state of the substrates, we performed molecular dynamics simulations on the complex model to simulate the dynamic conformational information of the superfamily enzymes and substrates during the enzymatic synthesis reaction process and obtained the simulation trajectory. Based on the simulation trajectory, we selected the optimal reaction conformation as the initial conformation. Multiple potential reaction pathways for enzyme-catalyzed synthesis are obtained based on the initial conformation; potential energy surface scanning, transition state structure optimization and single-point energy calculation are performed on each reaction pathway based on a combination of quantum mechanics and molecular mechanics to obtain the structural characteristics and energy data of reaction intermediates, transition states and products under each reaction pathway. Based on the enzymatic synthesis pathway of the N-nitrosourea group, a target reaction pathway was screened to identify the product structure that perfectly matches the characteristic structure of the N-nitrosourea group. Based on the structural characteristics and energy data of the target reaction pathway, the enzymatic synthesis mechanism of the N-nitrosourea group was quantitatively analyzed.

[0007] Preferably, determining the protonation state of the substrate includes: Obtain structural data and measured p of the superfamily enzymes and substrates that catalyze the synthesis of N-nitrosourea groups. K a Data and characteristics of the microenvironment of enzyme activity centers; Among them, the microenvironmental characteristics of the enzyme active site include residue type, hydrogen bond distance and hydrophobic interaction strength; Enzyme-substrate models were constructed using Autodock software, and the substrate models were converted into molecular graph features. These molecular graph features, along with the microenvironment features of the enzyme active site, were then input into a pre-trained GNN model to predict the p-values ​​of each ionizable site of the substrate. K a value; We constructed a computational model for quantum chemical methods and performed structural optimization using density functional theory and mixed basis sets in gas-phase and implicit solvent environments, respectively, to obtain optimized structural and energy data. Based on free energy calculations, the p-values ​​of each ionizable site of the substrate are obtained. K a The values ​​were compared with the predicted p values ​​for each ionizable site of the substrate. K a The value determines the protonation state of the substrate.

[0008] Preferably, determining the initial conformation includes: Using the high-resolution crystal structure of streptozotocin synthase as a template, the symmetrical chain, glycerol, and sulfate non-essential heteromolecules in the crystal were removed, and the substrate was added. N δ ,N ω -dihydroxy- N ω -Methyl-L-arginine and oxygen form an enzyme-Fe²⁺-O₂-substrate complex; hydrogen atoms are added using CHARMM's HBUILD module, and a TIP3P water sphere with a radius of 43 Å is constructed. Na⁺ is randomly added to neutralize the system charge and form a complex model. Based on the complex model, molecular dynamics simulations are performed to minimize energy, simulate heating and equilibrium of the complex, and obtain the simulated trajectory. Based on simulated trajectories, deep embedded clustering and a denoising autoencoder (DEC-DAE) are used to extract key reaction coordinate features from each conformation cluster obtained by DEC-DAE clustering, evaluate the reactivity of each conformation cluster, screen out the optimal reaction conformation, and use the optimal reaction conformation as the initial conformation.

[0009] Preferably, the acquisition of multiple potential reaction pathways for enzymatic synthesis includes: Determine the molecular structure of the enzyme-substrate complex and convert the molecular structure into the SMILES sequence; The SMILES sequences were input into the Transformer-MolGen generative model to obtain multiple potential response pathways; these pathways are as follows: Pathway I: Oxygen extracts hydrogen atoms from the substrate, and the oxygen atoms attack the substrate to form a bridging peroxy intermediate. The OO and NC bonds break, and the NN bonds form to obtain the product. Pathway II: Oxygen attacks the substrate to form a bridging intermediate, the OO and NC bonds break, the NN bonds form, and the product is obtained; Pathway III: The reaction intermediate undergoes intramolecular nucleophilic attack to form a three-membered ring intermediate, which then undergoes ring cleavage to yield the substrate; Pathway IV: O2 forms bidentate coordination with Fe²⁺ and attacks the substrate to obtain the product; Pathway V: Water molecules participate in hydrogen migration to obtain the product.

[0010] Preferably, obtaining the structural characteristics and energy data of reaction intermediates, transition states, and products along each reaction pathway includes: A combination of quantum mechanics and molecular mechanics was used to perform potential energy surface scanning, transition state structure optimization, and single-point energy calculation for each reaction pathway. In potential energy surface scanning, intermediates are obtained by optimizing the endpoint of PES scanning, local minima are searched using the L-BFGS algorithm, the highest point of the scanning potential energy curve is set as the transition state, and the structural features of the transition state are obtained. In the optimization of the transition state structure, the density functional theory method B3LYP is used to optimize the structure and obtain the optimized transition state structure. In single-point energy calculations, based on the optimized reactant, transition state, and product structures, single-point energy calculations are performed to obtain the entire catalytic reaction potential energy surface, and potential energy curves for each reaction pathway are plotted.

[0011] Preferably, the quantitative analysis of the enzymatic synthesis mechanism of the N-nitrosourea group includes: Obtain the bond lengths, bond angles, hydrogen bond networks, and energy data for reaction intermediates, transition states, and products along each reaction pathway; The bond lengths, bond angles, hydrogen bond networks, and energy data of the key structures in each reaction pathway are input into the pre-trained GNN model to quantitatively analyze the energy contribution of key amino acid residues in the reaction process, and obtain the reaction rate-determining step corresponding to the highest energy barrier and the corresponding reaction pathway with thermodynamically stable intermediates in the pathway. Based on the obtained reaction pathway, energy barriers, residue interactions, and reaction mechanisms were analyzed to quantify the enzymatic synthesis mechanism of N-nitrosourea groups.

[0012] This invention also provides a system for elucidating the enzymatic synthesis mechanism of N-nitrosourea compounds, comprising: The state determination module is used to acquire structural data of the superfamily enzymes and substrates used for the catalytic synthesis of N-nitrosourea groups; based on the molecular structural characteristics in the substrate structural data and the microenvironment characteristics of the enzyme active site of the superfamily enzyme, it predicts the p-value of each ionizable site of the substrate. K a Value; based on p of each ionizable site of the substrate K a The protonation state of the substrate was determined by using quantum chemical methods to optimize the structure and calculate the free energy in gas phase and implicit solvent environments, respectively. The conformation module is used to acquire protein crystal structure data of superfamily enzymes and construct complex models including superfamily enzymes and substrates based on the protein crystal structure data of superfamily enzymes and substrate structure data. Based on the protonation state of the substrate, molecular dynamics simulations are performed on the complex model to simulate the dynamic conformation information of superfamily enzymes and substrates during enzymatic synthesis reactions and obtain simulation trajectories. Based on the simulation trajectories, the optimal reaction conformation is selected as the initial conformation. The pathway verification module is used to obtain multiple potential reaction pathways for enzyme-catalyzed synthesis based on the initial conformation; based on a combination of quantum mechanics and molecular mechanics, it performs potential energy surface scanning, transition state structure optimization and single-point energy calculation on each reaction pathway to obtain the structural characteristics and energy data of reaction intermediates, transition states and products under each reaction pathway. The analysis module is used to screen target reaction pathways based on the enzymatic synthesis reaction pathway of N-nitrosourea group, identify the product structural features that completely match the characteristic structure of N-nitrosourea group, and quantitatively analyze the enzymatic synthesis mechanism of N-nitrosourea group based on the structural features and energy data of the target reaction pathway.

[0013] This invention also provides an electronic device, including a memory and a processor; The memory is used to store computer programs; When the processor executes the computer program stored in the memory, it implements the steps of the method described above for analyzing the enzymatic synthesis mechanism of N-nitrosourea compounds.

[0014] This invention also provides a computer-readable storage medium for storing a computer program, which, when executed by a processor, implements the steps of a method for analyzing the enzymatic synthesis mechanism of N-nitrosourea compounds as described above.

[0015] This invention provides a method and system for elucidating the enzymatic synthesis mechanism of N-nitrosourea compounds. Compared with the prior art, its advantages are as follows: This invention first predicts the p-values ​​of each ionizable site of the substrate. K a The substrate structure was optimized and free energy was calculated using quantum chemical methods. This process quantitatively determined the protonation state of the substrate using QM. Then, based on the protonation state of the substrate, molecular dynamics simulations were performed on the complex constructed from the enzyme and substrate to identify the optimal reaction conformation and use it as the initial conformation. Based on the initial conformation, multiple potential reaction pathways for enzyme-catalyzed reactions were obtained. This process was calculated using a combination of quantum mechanics and molecular mechanics to identify the structural characteristics and energy data of the reaction intermediates, transition states, and products under the most stable pathway at the atomic level. This was done to quantify the enzymatic synthesis mechanism of the N-nitrosourea group, thus establishing a complete quantitative relationship from the microenvironmental characteristics of the enzyme active site to the protonation state of the substrate and then to the reaction activity. This revealed the complete mechanism of the enzymatic synthesis of the N-nitrosourea pharmacophore at the atomic level. Attached Figure Description

[0016] Figure 1 A schematic diagram of the calculation process for a method to analyze the enzymatic synthesis mechanism of N-nitrosourea compounds provided in this embodiment of the invention; Figure 2 This is a schematic diagram of the potential energy profile of path I calculated at the B3LYP theoretical level for a method to analyze the enzymatic synthesis mechanism of N-nitrosourea compounds provided in this embodiment of the invention. Detailed Implementation

[0017] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0018] Currently, there is no publicly available complete analytical method for studying the atomic-level mechanism of enzyme-catalyzed N-nitrosourea formation. Experimental methods such as crystallography and kinetic analysis are insufficient to capture transient intermediates and transition states. Traditional computational methods such as quantum mechanics or molecular dynamics cannot accurately describe the reaction pathway because they cannot account for both electronic effects and the influence of the protein environment. Multi-scale computational processes rely on manual intervention (such as subjective screening of molecular dynamic conformations, empirical judgment of protonation states, and manual pre-setting of reaction pathways), resulting in the omission of potential reaction pathways (manually pre-set pathways easily miss unknown mechanisms) and insufficient quantification of the effects of key residues. At present, there is no transferable intelligent computational framework for other N-nitrosourea natural products, which makes it difficult to meet the need for efficient analysis of complex enzymatic reaction mechanisms. Therefore, developing an AI-driven multi-scale computational method for analyzing the enzymatic synthesis mechanism of N-nitrosourea pharmacophores is a technical challenge that urgently needs to be solved in this field.

[0019] To address the shortcomings of current technologies, this invention constructs a multi-scale computational framework that includes AI-assisted quantitative determination of protonation states, automatic conformation selection based on MD simulation, generative AI prediction with full-path verification using QM / MM, and quantitative analysis of GNN attribution mechanisms. Figure 1 As shown, this framework is a novel intelligent research and analysis method for studying the enzymatic synthesis mechanism of N-nitrosourea pharmacophores. Through AI-assisted determination of substrate quantification state, automatic screening of initial conformations based on MD simulation, and generative AI-based expansion of reaction pathways, it achieves, for the first time, atomic-level mechanistic analysis of this reaction. This provides a basis for developing new computational schemes for handling similar enzymatic reactions and also provides a theoretical foundation for efficient enzyme design and drug biosynthesis. Specifically, it includes the following steps: I. AI-assisted determination of the protonation state of the substrate.

[0020] GNN model predicts each ionizable site p of the substrate. K a Based on the enzyme's crystal structure data, a preliminary QM calculation model was constructed and p was performed. K a Numerical computation tasks; QM calculations are performed by the Gaussian program, and the results are used to determine the state of ionizable protons on the substrate.

[0021] II. Automatically select the initial structure based on MD simulation data.

[0022] Using crystal structure data of streptozotocin synthase, an active site model was constructed, including iron ions, substrate L-DNMA and O2, coordination residues, and water molecules. A complete system including a pre-equilibrium water sphere was established using CHARMM force field preprocessing. A series of energy minimizations were performed before MD simulation, followed by molecular dynamics simulation. An unsupervised clustering algorithm was used to automatically select the optimal structure as the initial conformation for subsequent calculations.

[0023] III. Generative AI prediction + QM / MM calculation to verify reaction pathways.

[0024] Generative AI was used to expand potential reaction pathways, and QM / MM calculations were performed to verify the feasibility of the pathways. The calculation model was divided into QM region and MM region. The QM region mainly involves molecules and amino acid residues involved in the reaction, while the rest is the MM region. QM / MM calculations were performed using the ChemShell software package. Based on the calculation results, the optimized stable configuration, potential energy surface scan curve, transition state structure, and single-point energy values ​​were obtained.

[0025] IV. GNN attribution analysis results determine the reaction mechanism.

[0026] Based on the calculation results, the structural and energy information of reactants, transition states, intermediates and products are analyzed. The role of key residues is quantified by the GNN model, the rate-determining step is automatically identified, and the mechanism conclusions with confidence and dynamic visualization results are output to determine the feasible reaction mechanism.

[0027] To further elaborate on the above scheme, it includes: I. AI-assisted determination of the protonation state of the substrate.

[0028] (1) GNN model construction and p K a predict.

[0029] Data foundation: Collect structural data of superfamily enzymes and substrates, and measure p K a Data and active site microenvironment characteristics (residue type, hydrogen bond distance, hydrophobic interaction strength) were used to construct an annotated dataset (input: substrate molecule map + active site microenvironment characteristics; output: p of each ionizable site). K a (Value). Model selection: SchNet graph neural network was adopted, which can directly process molecular topology (bond length, bond angle) and microenvironment features, and automatically extract substrate-protein interaction pairs. K a The influence weight.

[0030] Prediction process: ① Use Autodock to construct an initial model of the enzyme active site (containing Fe²⁺, 3 coordinating amino acids, substrate L-DNMA, and O2).

[0031] ② Transform the model into molecular graph features (nodes: atom type; edges: bond type and bond length) and microenvironment features (such as the hydrogen bond distance between R48 and the substrate, and the polar contribution of Q398), and input them into the pre-trained GNN model.

[0032] ③ The model outputs the p-values ​​of each ionizable site (e.g., hydroxyl, amino) of the substrate. K a Predicted values ​​(error < 0.3) and generated a "Microenvironmental Impact Report" (e.g., Y459 causes substrate amino groups to react via π-π interactions). K a (Reduced by 0.2 units).

[0033] (2) QM calculation verification p K a data.

[0034] Based on enzyme protein crystal structure data, substrate pontoonization and deprotonation models were constructed using GaussianView. The constructed computational models included iron ions, oxygen, the substrate, and iron-coordinated amino acid residues. Based on the calculated energy information, the gas-phase free energy and dissolution free energy were obtained. Substituting these into the pKa calculation formula, the pKa value was obtained. K a Numerical values, specifically including: In the QM model, the protonated substrate model and the deprotonated substrate model were geometrically optimized and frequency analysis was performed using the UB3LYP / 6-31G(d,p)&LAN2DZ level.

[0035] The input files for optimization calculations and frequency analysis in the gas phase are as follows: Calculate the input file: #p opt freq ub3lyp / 6-31G(d,p)&LANL2DZ.

[0036] For optimization calculations and frequency analysis in implicit solvent environments, the input file must include the implicit solvent model keyword SMD.

[0037] All structures were checked and found to have no imaginary frequencies.

[0038] The input file for higher-level single-point energy calculation is as follows: #p ub3lyp / 6-311++G(d,p)&LAN2TZ geom=allcheck.

[0039] The Gibbs free energy calculations for both gas-phase and implicit solvent environments were performed using the B3LYP density functional.

[0040] The input file is as follows: #p ub3lyp / 6-311++G(d,p)&LAN2TZ geom=allcheck.

[0041] The calculation of Gibbs free energy under implicit solvent conditions requires the inclusion of the implicit solvent model keyword SMD.

[0042] Data were collected to obtain the gas phase free energy and the solution free energy. The free energy of the solute under standard conditions (298.15 K, 1 M) in the solvent environment was obtained using the following formula: G aq = G g + G solv +1.89 kcal / mol p is calculated using the following formula K a The value is represented as: II. Automatically select the initial structure based on MD simulation data.

[0043] (1) MD simulation.

[0044] Based on data provided in the crystal database, a complete protein chain was extracted, and L-DNMA and oxygen were added; p was calculated using PROPKA. K a Numerical values ​​were used to determine the protonation state of all titratable amino acids; missing hydrogen atoms were added using the HBUILD module in CHARMM; the system was solvated using a water sphere, and sodium ions were randomly added to neutralize the system charge.

[0045] A series of energy minimizations were performed on the enzyme-substrate model using CHARMM. Subsequently, under random boundary conditions at 298 K, a kinetic simulation of the entire system was conducted using the CHARMM all-atom force field for 30 ns. The RMSD changes of the protein backbone atoms were monitored during the kinetic simulation. The system reached equilibrium after about 9 ns, with an RMSD value of approximately 1.9 Å.

[0046] (2) Unsupervised clustering to select the optimal initial conformation.

[0047] For the MD simulation trajectory (production stage: 298.15K, 30,000,000 steps), the DEC-DAE (Deep Embedded Clustering-Denoising Autoencoder) algorithm was used: ① Extract 12 key features: Fe-O2 bond length, Fe-L-DNMA coordination bond length, O1-C ω Distance and O1-H ω Distance, hydrogen bond distance between R48 / Q398 / Y459 and the substrate, etc.

[0048] ② Automatic clustering into 5-8 conformational clusters, and calculation of the reactivity score for each cluster (e.g., O1-C within the cluster). ω The distances are concentrated in the range of 3.3–3.5 Å and O1–H ωThe optimal range for distances between 2.5 and 3.0 Å is the highest, resulting in the best score.

[0049] ③ Output the top 3 cluster center conformations with the highest activity (such as the conformation with RMSD=1.9±0.1Å) as the initial structure for QM / MM calculation, which improves the screening accuracy by 50%.

[0050] III. Generative AI prediction + QM / MM calculation to verify reaction pathways.

[0051] (1) Generative AI predicts reaction pathways.

[0052] Data foundation: Collect a database of catalytic reactions of non-heme iron enzymes (including reactants, intermediates, transition states, product structures and reaction types), and construct an enzyme-substrate-reaction pathway sequenced dataset (molecular structures are converted into SMILES sequences).

[0053] Model selection: The Transformer-MolGen generative model was adopted (encoder: extracts enzyme-substrate complex features; decoder: generates SMILES sequences of potential reaction pathways).

[0054] Path generation: Input the SMILES sequence of the SznF-Fe²⁺-O2-L-DNMA complex, and the model generates 5 potential reaction pathways, including the original 3 pathways (I / II / III) and 2 new pathways (IV / V), and outputs the initial probability score of each pathway (based on the similarity to known enzymatic reactions, pathway I has the highest score, 0.85).

[0055] Pathway I: Oxygen first extracts hydrogen atoms from the substrate, then the OO and NC bonds break, and the NN bonds form, yielding the product.

[0056] Pathway II: Oxygen first attacks the substrate to form a bridging intermediate, and the subsequent steps are the same as in Pathway I.

[0057] Pathway III: The reaction intermediate undergoes intramolecular nucleophilic attack to form a three-membered ring intermediate, followed by ring cleavage to obtain the substrate.

[0058] Path IV: O2 first forms a bidentate coordination with Fe²⁺ before attacking the substrate.

[0059] Pathway V: Hydrogen migration pathway involving water molecules.

[0060] (2) QM / MM calculate possible reaction pathways.

[0061] QM / MM calculations were performed using the ChemShell software package. The Turboomole program was used to calculate the QM region, and the DL_POLY program was used to calculate the MM region. The QM / MM calculations included: potential energy surface (PES) scanning, optimization of reactant, transition state, and intermediate structures using the HDLC optimizer combined with the L-BFGS algorithm; and the use of B3LYP hybrid functionals, with the all-electron basis set Wachters+f for iron ions and the 6-31G(d,p) basis set for nonmetal atoms.

[0062] IV. GNN attribution analysis results determine the reaction mechanism.

[0063] Construct a structure-energy-action related GNN model, inputting the structural features (bond lengths, bond angles, hydrogen bond networks) and energy data (energy barriers, free energies) of reactants / transition states / intermediates calculated by QM / MM, and outputting: Automatic identification of rate-determining step: Based on energy barrier data, TS1 (energy barrier 20.8 kcal / mol) is labeled as the rate-determining step.

[0064] Quantification of the role of key residues: For example, R48 lowers the TS1 energy barrier by 1.2 kcal / mol through hydrogen bonding, Q398 stabilizes the IM2 intermediate through electrostatic interaction (contributing -0.8 kcal / mol free energy), and Y459 fixes the substrate conformation through hydrophobic interaction.

[0065] Path feasibility assessment: Paths II / III and the two new paths are automatically marked as "infeasible probability > 99%" due to the continuous increase in energy curves (no stable intermediates), while path I is marked as "feasible probability 95%".

[0066] Specific experiment: I. AI-assisted determination of the protonation state of the substrate.

[0067] (1) GNN model construction and p K a predict.

[0068] Dataset partitioning: Collect 100+ non-heme iron enzyme active center samples and partition them into training set, validation set and test set in a 7:2:1 ratio.

[0069] Training parameters: Optimizer AdamW, learning rate 1e-4, batch size 32, training epochs 100, early stopping strategy (stop if the validation set error does not decrease for 5 consecutive epochs).

[0070] Test set performance: p K a The predicted MAE (mean absolute error) is 0.28, and R² is 0.94.

[0071] (2) QM calculation verification pK a data.

[0072] The substrate protonation and deprotonation models were constructed using GaussianView. In the initial QM calculation model, the coordination bond length between iron ions and oxygen was 1.95 Å, and the coordination bond length between the substrate and iron ions was 2.06 Å. For all structures, density functional theory and mixed basis sets were used for structural optimization in gas phase and implicit solvent environments, respectively, to obtain optimized bond length data. The QM calculations were mainly performed using the Gaussian program.

[0073] In the optimized structure, the coordination bond length between iron ions and oxygen is 1.98 Å, and the coordination bond length between the substrate and iron ions is 2.00 Å. To obtain more accurate energies, single-point energy calculations and Gibbs free energy calculations were performed on a larger basis set based on the optimized structure.

[0074] Based on the obtained gas phase free energy and dissolution free energy, using the formula G aq and formula p K a p was calculated K a Value; Calculation result: p of three ionizable protons in L-DNMA K a The values ​​were 11.7, 15.0, and 27.3, respectively. At the experimental pH of 7.5, the substrate was in a protonated state, consistent with the GNN prediction.

[0075] Parameter selection criteria: The B3LYP functional combines accuracy and efficiency in handling the interaction between transition metals and organic molecules and has been verified to be suitable for the study of non-heme iron enzymes; the SMD implicit solvent model can simulate the internal polar environment of proteins, which is consistent with the microenvironment of the enzyme active site.

[0076] II. Automatically select the initial structure based on MD simulation data.

[0077] This invention uses the high-resolution crystal structure (6M9R) of streptozotocin synthase as a template; the crystal structure retains the complete ligand in the substrate pocket of the cupin region and has high resolution, which is beneficial for the precise positioning of metal ions, substrates and oxygen molecules.

[0078] (1) MD simulation.

[0079] Structural pretreatment: Removal of non-essential heterologous molecules such as symmetric chains, glycerol, and sulfate from the crystal; addition of substrate. N δ , N ω -dihydroxy- N ω-Methyl-L-arginine and oxygen were used to form an enzyme-Fe²⁺-O₂-substrate complex; hydrogen atoms were added using the HBUILD module of CHARMM to complete the complex, and a TIP3P water sphere with a radius of 43 Å was constructed; Na⁺ was randomly added to neutralize the system charge; the CHARMM all-atom force field was used; substrate missing parameters were obtained using SwissParam.

[0080] After completing the structural preprocessing and determining the force field parameters, molecular dynamics simulations are performed, mainly including energy minimization, heating, and equilibrium; specifically including: Energy minimization: After the system construction is completed, the energy minimization stage is entered: first, the steepest descent method is executed for 100 steps, and then the conjugate gradient method is added for 200 steps. During this period, a resonant constraint of 10 kcal·mol⁻¹·Å⁻² is applied to the heavy atoms of the protein backbone to eliminate van der Waals conflicts and high-energy contacts in the initial structure.

[0081] Heating and Equilibrium: Heating phase (gradually increasing the temperature from 0K to 298.15K over 200,000 steps (200ps)); Equilibrium phase (continuing to run at 298.15K for another 200,000 steps (200ps) to further relax the system); Production phase (running at 298.15K for 30,000,000 steps (30ns) for formal sampling).

[0082] (2) Unsupervised clustering to select the optimal initial conformation.

[0083] After the production phase, the RMSD of the protein backbone was calculated to assess structural stability; fluctuations in total energy, potential energy, and kinetic energy over time were extracted to confirm energy convergence; after 9 ns, the RMSD stabilized at 1.9 ± 0.1 Å, indicating that the system had reached full equilibrium; DEC-DAE clustering yielded 6 clusters, with the top cluster (activity score 0.92) showing O1-C ω The distance is concentrated at 3.32±0.05 Å, and it was selected as the initial conformation for QM / MM.

[0084] Parameter selection criteria: A simulation duration of 30 ns can ensure the capture of protein conformational equilibrium (preliminary experiments show that the RMSD fluctuates significantly within 9 ns); the TIP3P water model has good compatibility with the CHARMM force field and can accurately simulate solvent effects.

[0085] III. Generative AI prediction + QM / MM calculation to verify reaction pathways.

[0086] (1) Generative AI predicts reaction pathways.

[0087] Model training: 6 layers each for Transformer-MolGen encoder / decoder, 256 hidden layer dimensions, 8 attention heads, and training set loss reduced to 0.08.

[0088] Output paths: Among the 5 paths, the probability of path I (hydrogen migration-attack) is 0.85, path II (direct attack) is 0.12, path III (tripartite ring) is 0.02, new path IV (bident coordination) is 0.01, and new path V (water molecule participation) is 0.00.

[0089] Pathway I (Hydrogen Migration-Attack Mechanism): Hydrogen dioxide extracts hydrogen from the hydroxyl group of the substrate; hydrogen dioxide attacks the substrate, forming a bridged peroxy intermediate; the OO bond breaks to generate Fe³⁺-OH and NO•; Fe³⁺-OH extracts H₂ via a water bridge. ω NO • Attack N ω It forms N-nitrosourea.

[0090] (2) QM / MM calculate possible reaction pathways.

[0091] Computational platform: ChemShell integrates Turbomole and DL_POLY, using hydrogen-linked atom methods to handle covalent boundaries and an electron-intercalation scheme to handle QM / MM electrostatic coupling.

[0092] QM region definition: With the cupin active center as the core, the QM region contains: Fe(II) ions; first coordination layer: imidazole rings of side chains of H407, H409, and H448; second coordination layer: side chains of Q398, Y459, and R48 and water molecules; all atoms of the substrate; oxygen.

[0093] Calculation method and basis set: B3LYP functional, considering dispersion correction; Wachters+f all-electron basis set for metals; 6-31G(d,p) basis set for nonmetals.

[0094] Potential energy surface scanning strategy: The intermediate is obtained by optimizing the endpoint of the PES scan, and the local minimum is searched using the L-BFGS algorithm; the highest point of the scanned potential energy surface is set as the transition state, and further optimized using the P-RFO algorithm in HDLC. Step size is set to 0.05~0.08 Å; key bond lengths: O1-H, O1-O2, N ω -NO; After each step, use the HDLC optimizer + L-BFGS algorithm for local optimization.

[0095] ① Optimization of reactant structure.

[0096] The reactant structure was optimized by selecting conformations from molecular dynamics simulations and then using density functional theory (DFT) to optimize the geometry. In the optimized reactant structure, the coordination bond length between the central iron ion and the coordinating histidine residue is approximately 2 Å; the coordination bond length with the substrate is 2–2.5 Å; the dioxygen ion is terminally monodentately coordinated with the iron center; the distance between O2 and the central iron ion is 2.03 Å; and the O1–O2 bond length is 1.26 Å. The optimized reactant structure was further refined using DFT.

[0097] To obtain more accurate energy, we perform single-point energy calculations on a larger basis set based on the previous structural optimization.

[0098] ② Hydrogen extraction steps.

[0099] The potential energy curve of the key chemical bond change process was obtained by using PES scanning. By analyzing the possible structural changes in the key chemical bond changes, the energy of each structure step was calculated. The structure corresponding to the highest energy point on the potential energy curve was used as the initial transition state structure to complete the preliminary modeling. The P-RFO algorithm in HDLC was used to further optimize and obtain the stable transition state structure.

[0100] O in transition state 1 2- The distance between H is shortened to 1.21 Å, the distance between Fe and O2 is increased to 2.69 Å, and the distance between Fe and O1 is shortened to 1.88 Å.

[0101] The intermediate was obtained by PES scanning endpoint optimization. The local minimum was searched using the L-BFGS algorithm. The geometric configuration was optimized using the density functional theory method B3LYP to obtain the optimized intermediate structure. In the obtained stable intermediate structure 1, the distance between O2 and H was shortened to 0.91 Å, and the distance between Fe and O1 was 1.80 Å.

[0102] To obtain more accurate energy, we perform single-point energy calculations on a larger basis set based on the previous structural optimization.

[0103] ③ Formation of bridging peroxide intermediates.

[0104] The potential energy curve of the key chemical bond change process was obtained by using PES scanning. By analyzing the possible structural changes in the key chemical bond changes, the energy of each structure step was calculated. The structure corresponding to the highest energy point on the potential energy curve was used as the initial transition state structure to complete the preliminary modeling. The P-RFO algorithm in HDLC was used to further optimize and obtain the stable transition state structure.

[0105] With O1-C ω The distance is used as the reaction coordinate for scanning, and O1-C in transition state 2 ωThe distance between O2 and Fe decreased from 3.53 Å to 1.71 Å, while the distance between O2 and Fe increased to 3.63 Å.

[0106] The intermediate was obtained by optimizing the endpoint of the PES scan. The local minimum was searched using the L-BFGS algorithm, and the geometric configuration was optimized using the density functional theory method B3LYP to obtain the optimized intermediate structure. In the stable intermediate structure 2, O1-C ω The distance is 1.45 Å, and O2 points towards the center of iron at a distance of 2.73 Å.

[0107] To obtain more accurate energy, we perform single-point energy calculations on a larger basis set based on the previous structural optimization.

[0108] ④ The breaking of OO bonds and the formation of NO•.

[0109] The potential energy curve of the key chemical bond change process was obtained by using PES scanning. By analyzing the possible structural changes in the key chemical bond changes, the energy of each structure step was calculated. The structure corresponding to the highest energy point on the potential energy curve was used as the initial transition state structure to complete the preliminary modeling. The P-RFO algorithm in HDLC was used to further optimize and obtain the stable transition state structure.

[0110] PES scanning was performed using the O1-O2 distance as the reaction coordinate. In transition state 3, the O1-O2 bond length increased from 1.47 Å to 1.80 Å. ω The bond length increases to 1.61 Å.

[0111] The intermediate was obtained through PES scan endpoint optimization. The local minimum was searched using the L-BFGS algorithm, and the geometric configuration was optimized using the density functional theory method B3LYP to obtain the optimized intermediate structure. In the stable intermediate structure 3, the O1-O2 bond length separated to 2.89 Å, while the NC... ω The bond length increased to 2.93 Å, and spin population analysis showed that NO exhibited free radical characteristics.

[0112] To obtain more accurate energy, we perform single-point energy calculations on a larger basis set based on the previous structural optimization.

[0113] ⑤ NO• attacks the substrate to form N-nitrosourea groups.

[0114] The potential energy curve of the key chemical bond change process was obtained by using PES scanning. By analyzing the possible structural changes in the key chemical bond changes, the energy of each structure step was calculated. The structure corresponding to the highest energy point on the potential energy curve was used as the initial transition state structure to complete the preliminary modeling. The P-RFO algorithm in HDLC was used to further optimize and obtain the stable transition state structure.

[0115] Nitric oxide free radicals attack the N on the substrate ω Previously, substrate conformation fine-tuning enabled H ω Pointing to the Fe(III)-OH unit, FeIII-OH-H ω The distance between them is 1.84 Å, NO-N ω The distance is 2.95 Å, and the Fe-H transition state 6... ω Shortened to 1.34 Å, N ω -H ω The bond length increases to 1.15 Å, NO-N ω Shortened to 1.81 Å.

[0116] The intermediate was obtained through PES scan endpoint optimization, and the local minimum was searched using the L-BFGS algorithm. The geometric configuration was optimized using the density functional theory method B3LYP to obtain the optimized intermediate structure. In the obtained stable product structure, NN... ω The bond length is 1.37 Å, the Fe-H2O coordination bond length is 2.20 Å, and the NN=O bond angle of the product N-nitrosourea is 118°, which is consistent with the structural information of streptozotocin.

[0117] To obtain more accurate energies, single-point energy calculations were performed on a larger basis set based on the previous structural optimization, yielding the entire catalytic reaction potential energy surface, thus identifying the reaction pathway. Finally, the potential energy surface diagram was plotted, as shown below. Figure 2 As shown.

[0118] ⑥ Verify other paths.

[0119] The potential energy curve of the key chemical bond change process was obtained by using PES scanning. By analyzing the possible structural changes in the key chemical bond changes, the energy of each structure step was calculated. The structure corresponding to the highest energy point on the potential energy curve was used as the initial transition state structure to complete the preliminary modeling. The P-RFO algorithm in HDLC was used to further optimize and obtain the stable transition state structure.

[0120] With O1-C ω PES scanning was performed using the distance and O1-C distance as reaction coordinates. The energy curve continued to rise, and a stable intermediate structure could not be obtained, thus ruling out other pathways.

[0121] Verification basis: The product structure of path I (NN=O bond angle 118°) is consistent with the characteristic structure of streptozotocin detected by mass spectrometry.

[0122] IV. GNN attribution analysis results determine the reaction mechanism.

[0123] Attribution results include: Rate-determining step: The energy barrier for obtaining transition state 1 is 20.8 kcal / mol, contributing 45% of the total reaction energy barrier.

[0124] Residue effects: R48 (hydrogen bond contribution -1.2 kcal / mol) > Q398 (electrostatic contribution -0.8 kcal / mol) > Y459 (hydrophobic contribution -0.5 kcal / mol).

[0125] Reaction mechanism: Fe II –O2 ・ ⁻Removal of a hydrogen atom from the substrate → Fe III –OOH attacks C ω → The OO and NC bonds break to generate NO. ・ →Fe III -OH binds H ω →NO ・ Attack N ω It forms an N-nitrosourea pharmacophore.

[0126] This invention is the first to deeply integrate AI algorithms with multi-scale computing, revealing the complete mechanism of enzymatic synthesis of N-nitrosourea pharmacophores at the atomic level and establishing an AI-driven multi-scale computing framework. It is also the first to quantitatively demonstrate, through GNN+QM coupling, that substrate quantification is a prerequisite for reaction initiation, and that key residues (R48 / Q398 / Y459) enable O1-C... ω Maintaining an optimal reaction distance of 3.32 Å; locking the "hydrogen migration-attack" pathway and the rate-determining energy barrier of 20.8 kcal / mol; eliminating four invalid pathways through generative AI; providing a complete energy profile with confidence (95%) at the atomic scale for the first time, providing an intelligent paradigm for the study of superfamily enzyme catalytic mechanisms; establishing a quantitative correlation model of "microenvironment-protonation state-reactivity", filling the technical gap of "AI attribution" in non-heme iron enzyme catalysis.

[0127] The mechanisms elucidated in this invention can guide the rational modification of industrial strains for anticancer drugs such as streptozotocin and chlorpromazine: based on GNN-quantified residue interactions, hydrogen bond networks can be optimized through targeted mutations; it can serve as the underlying mechanism module of an AI drug discovery platform: generative AI can rapidly expand the biosynthetic pathways of novel N-nitrosourea compounds, assisting in the green manufacturing of lead compounds for anticancer / hypoglycemic drugs; through transfer learning, it can quickly adapt to other non-heme iron enzymes (such as chlorpromazine synthase and plant N-nitroso compound synthase), promoting the intelligent upgrade of green biomanufacturing of natural products.

[0128] This invention is a transferable theoretical calculation implementation technology based on a molecular dynamics simulation (MD) and quantum mechanics / molecular mechanics (QM / MM) computational coupling framework for analyzing the formation mechanism of N-nitrosourea pharmacophore catalyzed by streptozotocin synthase. By automating, refining, and increasing the efficiency of multi-scale computational processes driven by AI, it solves the problems of reliance on manual labor, low efficiency, and missed path identification in traditional calculation methods, providing an intelligent multi-scale computational method for the study of the enzymatic synthesis mechanism of N-nitrosourea pharmacophore and related drug biosynthesis.

[0129] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.

Claims

1. A method for elucidating the enzymatic synthesis mechanism of N-nitrosourea compounds, characterized in that, Includes the following steps: Obtain structural data of the superfamily enzymes and substrates used for the catalytic synthesis of N-nitrosourea groups; based on the molecular structural features in the substrate structural data and the microenvironment characteristics of the enzyme active site of the superfamily enzymes, predict the p-values ​​of each ionizable site of the substrate. K a Value; based on p of each ionizable site of the substrate K a The protonation state of the substrate was determined by using quantum chemical methods to optimize the structure and calculate the free energy in gas phase and implicit solvent environments, respectively. We obtained the protein crystal structure data of the superfamily enzymes and constructed a complex model including the superfamily enzymes and substrates based on the protein crystal structure data of the superfamily enzymes and substrates. Based on the protonation state of the substrates, we performed molecular dynamics simulations on the complex model to simulate the dynamic conformational information of the superfamily enzymes and substrates during the enzymatic synthesis reaction process and obtained the simulation trajectory. Based on the simulation trajectory, we selected the optimal reaction conformation as the initial conformation. Multiple potential reaction pathways for enzyme-catalyzed synthesis are obtained based on the initial conformation; potential energy surface scanning, transition state structure optimization and single-point energy calculation are performed on each reaction pathway based on a combination of quantum mechanics and molecular mechanics to obtain the structural characteristics and energy data of reaction intermediates, transition states and products under each reaction pathway. Based on the enzymatic synthesis pathway of the N-nitrosourea group, a target reaction pathway was screened to identify the product structure that perfectly matches the characteristic structure of the N-nitrosourea group. Based on the structural characteristics and energy data of the target reaction pathway, the enzymatic synthesis mechanism of the N-nitrosourea group was quantitatively analyzed.

2. The method for elucidating the enzymatic synthesis mechanism of N-nitrosourea compounds according to claim 1, characterized in that, Determining the protonation state of the substrate includes: Obtain structural data and measured p of the superfamily enzymes and substrates that catalyze the synthesis of N-nitrosourea groups. K a Data and characteristics of the microenvironment of enzyme activity centers; Among them, the microenvironmental characteristics of the enzyme active site include residue type, hydrogen bond distance and hydrophobic interaction strength; Enzyme-substrate models were constructed using Autodock software, and the substrate models were converted into molecular graph features. These molecular graph features, along with the microenvironment features of the enzyme active site, were then input into a pre-trained GNN model to predict the p-values ​​of each ionizable site of the substrate. K a value; We constructed a computational model for quantum chemical methods and performed structural optimization using density functional theory and mixed basis sets in gas-phase and implicit solvent environments, respectively, to obtain optimized structural and energy data. Based on free energy calculations, the p-values ​​of each ionizable site of the substrate are obtained. K a The values ​​were compared with the predicted p values ​​for each ionizable site of the substrate. K a The value determines the protonation state of the substrate.

3. The method for elucidating the enzymatic synthesis mechanism of N-nitrosourea compounds according to claim 1, characterized in that, The determination of the initial conformation includes: Using the high-resolution crystal structure of streptozotocin synthase as a template, the symmetrical chain, glycerol, and sulfate non-essential heteromolecules in the crystal were removed, and the substrate was added. N δ ,N ω -dihydroxy- N ω -Methyl-L-arginine and oxygen form an enzyme-Fe²⁺-O₂-substrate complex; hydrogen atoms are added using CHARMM's HBUILD module, and a TIP3P water sphere with a radius of 43 Å is constructed. Na⁺ is randomly added to neutralize the system charge and form a complex model. Based on the complex model, molecular dynamics simulations are performed to minimize energy, simulate heating and equilibrium of the complex, and obtain the simulated trajectory. Based on simulated trajectories, deep embedded clustering and a denoising autoencoder (DEC-DAE) are used to extract key reaction coordinate features from each conformation cluster obtained by DEC-DAE clustering, evaluate the reactivity of each conformation cluster, screen out the optimal reaction conformation, and use the optimal reaction conformation as the initial conformation.

4. The method for elucidating the enzymatic synthesis mechanism of N-nitrosourea compounds according to claim 3, characterized in that, The acquisition of multiple potential reaction pathways for enzymatic synthesis includes: Determine the molecular structure of the enzyme-substrate complex and convert the molecular structure into the SMILES sequence; The SMILES sequences were input into the Transformer-MolGen generative model to obtain multiple potential response pathways; these pathways are as follows: Pathway I: Oxygen extracts hydrogen atoms from the substrate, and the oxygen atoms attack the substrate to form a bridging peroxy intermediate. The OO and NC bonds break, and the NN bonds form to obtain the product. Pathway II: Oxygen attacks the substrate to form a bridging intermediate, the OO and NC bonds break, the NN bonds form, and the product is obtained; Pathway III: The reaction intermediate undergoes intramolecular nucleophilic attack to form a three-membered ring intermediate, which then undergoes ring cleavage to yield the substrate; Pathway IV: O2 forms bidentate coordination with Fe²⁺ and attacks the substrate to obtain the product; Pathway V: Water molecules participate in hydrogen migration to obtain the product.

5. The method for elucidating the enzymatic synthesis mechanism of N-nitrosourea compounds according to claim 4, characterized in that, The acquisition of structural characteristics and energy data of reaction intermediates, transition states, and products along each reaction pathway includes: A combination of quantum mechanics and molecular mechanics was used to perform potential energy surface scanning, transition state structure optimization, and single-point energy calculation for each reaction pathway. In potential energy surface scanning, intermediates are obtained by optimizing the endpoint of PES scanning, local minima are searched using the L-BFGS algorithm, the highest point of the scanning potential energy curve is set as the transition state, and the structural features of the transition state are obtained. In the optimization of transition state structure, density functional theory is used to optimize the structure and obtain the optimized transition state structure. In single-point energy calculations, based on the optimized reactant, transition state, and product structures, single-point energy calculations are performed to obtain the entire catalytic reaction potential energy surface, and potential energy curves for each reaction pathway are plotted.

6. The method for elucidating the enzymatic synthesis mechanism of N-nitrosourea compounds according to claim 5, characterized in that, The quantitative analysis of the enzymatic synthesis mechanism of the N-nitrosourea group includes: Obtain the bond lengths, bond angles, hydrogen bond networks, and energy data for reaction intermediates, transition states, and products along each reaction pathway; The bond lengths, bond angles, hydrogen bond networks, and energy data of the key structures in each reaction pathway are input into the pre-trained GNN model to quantitatively analyze the energy contribution of key amino acid residues in the reaction process, and obtain the reaction rate-determining step corresponding to the highest energy barrier and the corresponding reaction pathway with thermodynamically stable intermediates in the pathway. Based on the obtained reaction pathway, energy barriers, residue interactions, and reaction mechanisms were analyzed to quantify the enzymatic synthesis mechanism of N-nitrosourea groups.

7. A system for elucidating the enzymatic synthesis mechanism of N-nitrosourea compounds, characterized in that, include: The state determination module is used to acquire structural data of the superfamily enzymes and substrates used for the catalytic synthesis of N-nitrosourea groups; Based on the molecular structural features in the substrate structure data and the microenvironment characteristics of the enzyme active site of the superfamily enzyme, predict the p-values ​​of each ionizable site of the substrate. K a Value; based on p of each ionizable site of the substrate K a The protonation state of the substrate was determined by using quantum chemical methods to optimize the structure and calculate the free energy in gas phase and implicit solvent environments, respectively. The conformation module is used to acquire protein crystal structure data of superfamily enzymes and construct complex models including superfamily enzymes and substrates based on the protein crystal structure data of superfamily enzymes and substrate structure data. Based on the protonation state of the substrate, molecular dynamics simulations are performed on the complex model to simulate the dynamic conformation information of superfamily enzymes and substrates during enzymatic synthesis reactions and obtain simulation trajectories. Based on the simulation trajectories, the optimal reaction conformation is selected as the initial conformation. The pathway verification module is used to obtain multiple potential reaction pathways for enzyme-catalyzed synthesis based on the initial conformation; based on a combination of quantum mechanics and molecular mechanics, it performs potential energy surface scanning, transition state structure optimization and single-point energy calculation on each reaction pathway to obtain the structural characteristics and energy data of reaction intermediates, transition states and products under each reaction pathway. The analysis module is used to screen target reaction pathways based on the enzymatic synthesis reaction pathway of N-nitrosourea group, identify the product structural features that completely match the characteristic structure of N-nitrosourea group, and quantitatively analyze the enzymatic synthesis mechanism of N-nitrosourea group based on the structural features and energy data of the target reaction pathway.

8. An electronic device, characterized in that, include: Memory and processor; The memory is used to store computer programs; When the processor executes the computer program stored in the memory, it implements the steps of the method for analyzing the enzymatic synthesis mechanism of N-nitrosourea compounds as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, Used to store a computer program, which, when executed by a processor, implements the steps of a method for analyzing the enzymatic synthesis mechanism of N-nitrosourea compounds as described in any one of claims 1 to 6.