Covalent molecule simulation method and device based on AMBER and computer readable storage medium

Through the streamlined and automated AMBER molecular simulation method and device, the problems of low efficiency and poor accuracy of manual operation have been solved, and efficient and reliable molecular dynamics simulation has been achieved to support drug development and material design.

CN120673870APending Publication Date: 2025-09-19KEYING FUTURE (SHANGHAI) INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510958074.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing molecular dynamics simulations suffer from low manual operation efficiency, poor accuracy and reliability, and difficulty in ensuring consistency and repeatability, which affects scientific research and industrial development.

Method used

Provided is an AMBER-based covalent molecular simulation method and device, which streamlines and automates the molecular dynamics simulation process, including protein structure processing, initial simulation system construction, simulation operation, data topology configuration and covalent molecular information generation, reducing manual operations and improving simulation accuracy and consistency.

Benefits of technology

It improves the efficiency of molecular dynamics simulation and the reliability of simulation results, meets the actual needs of enterprises, and supports the research process in fields such as drug development and material design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673870A_ABST
    Figure CN120673870A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of molecule simulation, and particularly discloses an AMBER-based covalent molecule simulation method and device and a computer readable storage medium, and the method comprises the following steps: obtaining a protein structure; constructing an initialization simulation system based on AMBER and the protein structure; executing simulation operation based on the initialized simulation system to obtain simulation operation data; performing topological configuration on the simulation operation data to generate configured data; and generating covalent molecule information based on the configured data. Molecular dynamics simulation is streamlined and automated, so that various defects of an existing manual operation method are effectively overcome, the working efficiency is improved, the simulation accuracy is improved, and the actual requirements of enterprises are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of molecular simulation, and in particular to an AMBER-based covalent molecular simulation method, device, and computer-readable storage medium. Background Art

[0002] Molecular dynamics simulation, as an important means to study the structure, dynamics and thermodynamic properties of molecules, plays a key role in many fields such as drug development, materials science, and biochemistry.

[0003] Research in molecular dynamics simulations has made considerable progress, including molecular dynamics simulations based on classical mechanics and ab initio molecular dynamics simulations based on quantum mechanics. Classical mechanics simulations offer high computational efficiency and are suitable for simulations of larger molecular systems and longer timescales, but they have limitations when dealing with processes involving changes in electronic structure. Quantum mechanics simulations can accurately describe the electronic structure and chemical bond properties of molecules, but they are computationally expensive, making them difficult to apply to large-scale molecular systems and long-duration simulations.

[0004] However, in practical applications, most relevant processes are still performed manually. Manual operations are not only inefficient and unable to meet the rapidly evolving needs of scientific research and industry, but are also prone to human error, which affects the accuracy and reliability of simulation results. Furthermore, manual operations cannot guarantee consistency and repeatability, and different researchers may obtain different results when performing the same simulation. This severely restricts the efficient application and development of molecular dynamics simulations in various fields. Summary of the Invention

[0005] In order to overcome the above-mentioned technical problems existing in the prior art, the embodiments of the present invention provide a covalent molecular simulation method, device and computer-readable storage medium based on AMBER. By streamlining and automating molecular dynamics simulation, the various defects of the existing manual operation methods are effectively overcome, work efficiency is improved, simulation accuracy is improved, and the actual needs of enterprises are met.

[0006] To achieve the above-mentioned objectives, an embodiment of the present invention provides a covalent molecular simulation method based on AMBER, the method comprising: obtaining a protein structure; constructing an initialization simulation system based on AMBER and the protein structure; performing a simulation run based on the initialization simulation system to obtain simulation run data; topologically configuring the simulation run data to generate configured data; and generating covalent molecular information based on the configured data.

[0007] Preferably, constructing an initialization simulation system based on AMBER and the protein structure includes: preprocessing the protein structure to obtain a preprocessed structure; inputting the preprocessed structure into AMBER, loading corresponding force field data, performing charge neutralization processing, and obtaining a first initial system; performing solvation processing on the first initial system based on a preset water molecule model to obtain a processed system; and configuring ion concentration on the processed system to obtain an initialization simulation system.

[0008] Preferably, the preprocessing of the protein structure to obtain a preprocessed structure includes: performing redundancy analysis on the protein structure to obtain redundant chains / heteroatoms; processing the protein structure based on the redundant chains / heteroatoms and a preset crystal water threshold to obtain a first processed structure; obtaining a homology model, and protonating the first processed structure based on the homology model to obtain a second processed structure; performing residue atom analysis on the second processed structure to obtain unidentified residue atoms; and processing the second processed structure based on the unidentified residue atoms to generate a preprocessed structure.

[0009] Preferably, the simulation operation is performed based on the initialization simulation system to obtain simulation operation data, including: performing energy minimization processing on the initialization simulation system to obtain a processing or system; performing simulated heating processing based on the treated system to obtain a heated system; performing balancing processing on the heated system to obtain a balanced system; performing production simulation based on the balanced system to obtain simulation operation data.

[0010] Preferably, the topological configuration of the simulation operation data to generate configured data includes: extracting non-covalent ligands and covalent ligands in the simulation operation data; performing parameterization processing on the non-covalent ligands to obtain first processed data; performing parameterization processing on the covalent ligands to obtain second processed data; and generating configured data based on the first processed data and the second processed data.

[0011] Preferably, the parameterized processing of the non-covalent ligand to obtain first processed data includes: obtaining all first non-ligand data in the non-covalent ligand; deleting the first non-ligand data from the non-covalent ligand to obtain first deleted data; performing proton ligand perfection processing on the first deleted data to obtain perfected data; performing normative processing on the perfected data to generate first processed data.

[0012] Preferably, the performing parameterized processing on the covalent ligand to obtain second processed data includes: obtaining second non-ligand data from the covalent ligand based on the covalent bond; deleting the second non-ligand data from the covalent ligand to obtain second deleted data; performing protonation processing on the second deleted data to obtain protonated data; performing connection chain processing on the protonated data to obtain connected data; performing charge distribution analysis on the connected data to obtain charge data; and generating second processed data based on the connected data and the charge data.

[0013] Preferably, the generating of covalent molecule information based on the post-configuration data includes: performing maximum common substructure analysis on the post-configuration data to generate analysis data; performing atomic transformation analysis on the analysis data based on the energy conservation rule to generate closed loop information; calculating and generating an original binding free energy difference based on the closed loop information; correcting the original binding free energy difference to generate a corrected energy difference; determining the ligand binding free energy based on the corrected energy difference; and generating covalent molecule information based on the ligand binding free energy.

[0014] Correspondingly, the present invention also provides a covalent molecular simulation device based on AMBER, which is applied to the method according to an embodiment of the present invention, and the device includes: a structure acquisition unit for acquiring a protein structure; a system construction unit for constructing an initialization simulation system based on AMBER and the protein structure; a simulation operation unit for performing a simulation operation based on the initialization simulation system to obtain simulation operation data; a configuration unit for topologically configuring the simulation operation data to generate configured data; and an information generation unit for generating covalent molecular information based on the configured data.

[0015] On the other hand, the present invention further provides a computer-readable storage medium having a computer program stored thereon, which implements the method provided by an embodiment of the present invention when the program is executed by a processor.

[0016] The technical solution provided by the present invention has at least the following technical effects:

[0017] By streamlining and streamlining the existing molecular dynamics simulation process and automating it with existing tools, we can effectively improve the work efficiency and simulation accuracy of technicians, enhance the reliability of simulation results, and meet the actual needs of enterprises.

[0018] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the following detailed description, they are used to explain the embodiments of the present invention, but do not constitute a limitation of the embodiments of the present invention. In the accompanying drawings:

[0020] Figure 1 This is a specific implementation flow chart of the AMBER-based covalent molecular simulation method provided by an embodiment of the present invention;

[0021] Figure 2 Schematic diagram of the structure of the AMBER-based covalent molecular simulation device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0022] The following describes the specific implementation of the embodiment of the present invention in detail with reference to the accompanying drawings. It should be understood that the specific implementation described herein is only used to illustrate and explain the embodiment of the present invention and is not used to limit the embodiment of the present invention.

[0023] The terms "system" and "network" in the embodiments of the present invention can be used interchangeably. "Multiple" refers to two or more. In view of this, "multiple" can also be understood as "at least two" in the embodiments of the present invention. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / ", unless otherwise specified, generally indicates that the previous and next associated objects are in an "or" relationship. In addition, it should be understood that in the description of the embodiments of the present invention, words such as "first" and "second" are only used to distinguish the purpose of description, and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying order.

[0024] See Figure 1 , an embodiment of the present invention provides a covalent molecular simulation method based on AMBER, the method comprising:

[0025] S10: Obtain protein structure;

[0026] S20: constructing an initial simulation system based on AMBER and the protein structure;

[0027] S30: Execute simulation operation based on the initialized simulation system to obtain simulation operation data;

[0028] S40: Performing topological configuration on the simulation operation data to generate configured data;

[0029] S50: Generate covalent molecule information based on the configured data.

[0030] In one possible embodiment, protein structure information is first obtained. For example, protein structures related to the target research can be retrieved from the PDB database to obtain multiple homologous protein structures. In actual application, since the protein structure information obtained from the PDB database is large and relatively messy, it is possible to select the protein structure information based on the resolution (preferably higher than the resolution). ), ligand binding state (selecting structures with similar binding states to the target ligand), and ultimately selecting a suitable protein structure (e.g., the protein structure with ID 1ABC). For regions of the structure with lower resolution, structural refinement is performed using cryo-electron microscopy (Cryo-EM) data and related software to improve structural accuracy. AMBER (a commonly used chemical molecular dynamics analysis software) is then used for analysis. Initial simulations are first constructed based on the protein structure file.

[0031] In an embodiment of the present invention, the initialization simulation system is constructed based on AMBER and the protein structure, including: preprocessing the protein structure to obtain a preprocessed structure; inputting the preprocessed structure into AMBER, loading the corresponding force field data, performing charge neutralization processing, and obtaining a first initial system; performing solvation processing on the first initial system based on a preset water molecule model to obtain a processed system; and configuring the ion concentration of the processed system to obtain an initialization simulation system.

[0032] In one possible embodiment, since the initially obtained protein structure information data has defects, it is necessary to preprocess it to ensure the accuracy of subsequent analysis. In an embodiment of the present invention, the preprocessing of the protein structure to obtain a preprocessed structure includes: performing redundancy analysis on the protein structure to obtain redundant chains / heteroatoms; processing the protein structure based on the redundant chains / heteroatoms and a preset crystal water threshold to obtain a first processed structure; obtaining a homology model, and protonating the first processed structure based on the homology model to obtain a second processed structure; performing residue atom analysis on the second processed structure to obtain unidentified residue atoms; processing the second processed structure based on the unidentified residue atoms to generate a preprocessed structure.

[0033] Specifically, we first use pdb4amber to perform redundancy analysis on the protein structure and analyze the redundant chains / heteroatoms therein, and then identify the redundant chains / heteroatoms based on the preset crystal water threshold (e.g., the distance from the protein atom is more than The protein structure is processed by removing redundant and unnecessary parts (such as crystal water that does not participate in key interactions) to obtain the first processed structure. Then, a homology model is obtained (for example, obtained by homology modeling) or a loop is constructed from scratch to simulate the missing region. Using PDB2PQR and PropKa software to assist, the protonated residues such as ARG, LYS, ASP, GLU and the protonation state of HIS residues are considered to protonate the first processed structure to obtain the second processed structure. Finally, the second processed structure is subjected to residue atomic analysis to find unidentified residue atoms, and the second processed structure is processed based on these atoms to generate the pre-processed structure. Finally, the output is a PDB file that conforms to the AMBER naming rules.

[0034] In the embodiment of the present invention, the pre-processed structure is then input into the AMBER20 software, and the corresponding ff19SB / GAFF2 force field data is loaded. ff19SB is commonly used for protein simulation, while GAFF2 is suitable for small molecule ligands. Reasonable selection of force field can improve the accuracy of the simulation, and then charge neutralization treatment is performed (such as adding Na + / Cl-) to obtain the first initial system. By neutralizing the charge, the entire molecular system is made electrically neutral, consistent with the actual physical and chemical environment. Otherwise, the charge imbalance will lead to unreasonable electrostatic interactions during the simulation, affecting the simulation results.

[0035] Then, the first initial system is solvated based on the TIP3P water molecule model, such as placing the system in a reasonable solvent environment (TIP3P water box, ), the protein is placed in a simulated solvent environment. TIP3P is a commonly used water molecule model. Setting an appropriate boundary distance can ensure that the protein has sufficient activity space during the simulation process and avoid the influence of boundary effects, thereby obtaining a treated system. The treated system is configured for ion concentration, and ions of appropriate concentration (such as 0.15M NaCl) are added to simulate the physiological salt concentration environment, further making the simulation system closer to the real biological environment, and obtaining an initialized simulation system. At this point, the initialization simulation system is constructed, and the topology (.prmtop) and coordinate files (.inpcrd) are saved. These two files contain the structure and parameter information of the simulation system, laying a solid foundation for subsequent simulation runs. At this point, the corresponding simulation run is executed.

[0036] In an embodiment of the present invention, performing simulation operation based on the initialized simulation system to obtain simulation operation data includes: performing energy minimization processing on the initialized simulation system to obtain a processing or system; performing simulated heating processing based on the treated system to obtain a heated system; performing balancing processing on the heated system to obtain a balanced system; performing production simulation based on the balanced system to obtain simulation operation data.

[0037] In one possible implementation, the initial simulation system is first minimized in steps using the steepest descent method and the conjugate gradient method to eliminate unreasonable interatomic distances and interactions in the system, bringing the system to a stable, lower energy state. For example, the steepest descent method is used for 500 iterations to rapidly reduce the system energy and eliminate large interatomic conflicts. The conjugate gradient method is then used for 1000 iterations to achieve a more stable energy state, resulting in the final system. In the initially constructed system, the atomic positions may be conflicting or improperly arranged. The energy minimization algorithm adjusts the atomic positions to reduce the system energy, laying the foundation for subsequent simulations.

[0038] Then, a simulated heating treatment is performed on the treated system to obtain a heated system. Specifically, the system temperature is gradually raised to the target temperature in the temperature range of 0-300K to simulate the process of the system from low temperature to physiological temperature. For example, starting from 0K, the temperature is gradually increased slowly at a rate of 50K per step. After each heating step, a 100ps equilibrium simulation is performed to observe the system structure and energy changes to ensure that the system smoothly transitions to 300K. In this process, the thermal motion of atoms gradually increases, and the energy and structure of the system also change accordingly. Then, the heated system is further subjected to equilibrium treatment. Specifically, equilibrium is performed under NVT (constant temperature and constant volume) or NPT (constant temperature and constant pressure) conditions. NVT conditions keep the temperature and volume of the system constant, while NPT conditions keep the temperature and pressure constant. For example, running under the NPT ensemble for 500ps allows the system to reach equilibrium under temperature (300K) and pressure (1atm) conditions to obtain the equilibrium system. The equilibrium process is to allow the system to reach a stable state under specific conditions, so that the conformation and dynamic behavior of the molecules no longer change significantly over time. At this time, the various physical and chemical properties of the system tend to be stable, providing a reliable initial state for production simulation.

[0039] Finally, a production simulation (ns-μs level) is performed on the equilibrated system. During this stage, a large amount of simulation data is acquired for analyzing molecular structural changes and dynamic behavior. The simulation duration is set based on actual needs. A longer simulation time can capture slow molecular processes and rare events, such as protein conformational transitions. For example, using MPI parallel computing technology, the simulation task is distributed to eight computing nodes for simultaneous execution. The simulation duration is set to 10ns, and the simulation trajectory and intermediate data are saved every 100ps to obtain the simulation data. At this point, the topology is configured based on the simulation data to generate the configured data.

[0040] In an embodiment of the present invention, the topological configuration of the simulation operation data to generate configured data includes: extracting non-covalent ligands and covalent ligands in the simulation operation data; performing parameterization processing on the non-covalent ligands to obtain first processed data; performing parameterization processing on the covalent ligands to obtain second processed data; and generating configured data based on the first processed data and the second processed data.

[0041] In one possible implementation, non-covalent ligands and covalent ligands are first extracted from the simulation run data and parameterized. In this embodiment of the present invention, parameterizing the non-covalent ligands to obtain first processed data includes: obtaining all first non-ligand data from the non-covalent ligands; deleting the first non-ligand data from the non-covalent ligands to obtain first deleted data; performing proton ligand perfection processing on the first deleted data to obtain perfected data; and performing standardization processing on the perfected data to generate first processed data.

[0042] Specifically, first obtain the first non-ligand data in the non-covalent ligand, specifically all atoms / chains except the ligand (such as solvent molecules, impurities, etc.), and then separate the ligand from the complex protein-ligand system so that the ligand can be parameterized separately. Then perform proton ligand perfection processing on the first deleted data to supplement the missing protons. Specifically, the ligand is protonated and checked for correctness to ensure that the protonation state of the ligand is consistent with the actual situation, because the protonation state will affect the charge distribution of the ligand and the interaction with the protein. Finally, perform normative processing on the improved data, such as making all atom names unique and outputting a mol2 file to avoid parameter errors caused by repeated atom names. This file contains the structure and partial chemical information of the ligand and is the basis for subsequent acquisition of charge and generation of parameters; and use antechamber to obtain charges and calculate the charge distribution of the ligand to make the charge of the ligand more consistent with the actual situation. Then, use parmchk2 to generate missing parameters. Parmchk2 automatically generates missing force field parameter files (.frcmod) based on the ligand's structural information. This file complements the ligand's force field information and ensures the ligand interacts correctly with the protein during the simulation. Finally, use parmed to verify the bond length and angle parameters and ensure their rationality and accuracy. Finally, generate the corresponding first processed data.

[0043] In an embodiment of the present invention, performing parameterized processing on the covalent ligand to obtain second processed data includes: obtaining second non-ligand data from the covalent ligand based on a covalent bond; deleting the second non-ligand data from the covalent ligand to obtain second deleted data; performing protonation processing on the second deleted data to obtain protonated data; performing connection chain processing on the protonated data to obtain connected data; performing charge distribution analysis on the connected data to obtain charge data; and generating second processed data based on the connected data and the charge data.

[0044] In one possible embodiment, all atoms / chains except the ligand + CYS are deleted from the covalent ligand, retaining the parts related to the formation of the covalent bond, i.e., the ligand and the CYS residues involved in the covalent binding. The second deletion or data is then protonated, and the chains are connected to form the same residue name, so that the ligand and CYS residues form a structural whole, which is convenient for subsequent parameterization. Finally, the charge distribution analysis is performed on the connected data. Specifically, the charges of the ligand and protein atoms (antechamber) are obtained and merged separately. The charges of the ligand and protein atoms are obtained separately by different calculation methods, such as using the AM1-BCC method to calculate the charge, and then they are merged to accurately describe the charge distribution of the covalent system. Add a connection to the protein to clarify the connection mode and position of the covalent bond. Due to the particularity of the covalent bond, the parameterization results also need to be manually verified for the covalent bond connection to ensure the stability of the covalent bond during the simulation process. Finally, the covalent bond stability is tested by short-term MD to verify whether the covalent bond can remain stable in the actual simulation environment to avoid unreasonable situations such as covalent bond breakage. Finally, the corresponding second processed data is generated, that is, the configured data is obtained. At this time, the processed non-covalent ligand data and the covalent ligand data can be summarized to prepare for the subsequent generation of comprehensive and accurate configured data.

[0045] In an embodiment of the present invention, the generating of covalent molecule information based on the configured data includes: performing maximum common substructure analysis on the configured data to generate analysis data; performing atomic transformation analysis on the analysis data based on the energy conservation rule to generate closed loop information; calculating and generating an original binding free energy difference based on the closed loop information; correcting the original binding free energy difference to generate a corrected energy difference; determining the ligand binding free energy based on the corrected energy difference; and generating covalent molecule information based on the ligand binding free energy.

[0046] In one possible implementation, the maximum common substructure of the configured data is first analyzed to generate analysis data. Specifically, the maximum common substructure between the ligands is found through analysis, and the relevant atoms are marked, such as the common skeleton (black), charged atoms (red), and discharged atoms (blue), to help understand the key parts and charge distribution of the protein-ligand interaction, and provide a basis for the subsequent creation of TI pairs and construction of closed loops. Then, according to the principle of conservation of energy, the analysis data is subjected to atomic transformation analysis to generate closed loop information. Specifically, LOMAP and FESetup can be used to create TI pairs and construct closed loops. LOMAP is used to find the mapping relationship between the protein and the ligand and determine which atoms need to be transformed during the simulation process; FESetup is used to set the simulation topology file and coordinate file and prepare the input data required for the simulation. By constructing a closed loop, the energy conservation of the simulation process can be ensured and the accuracy of the binding free energy calculation can be improved.

[0047] At this point, the original binding free energy difference is further calculated based on the closed loop information. For example, tools such as AmberTools and alchemlyb can be used to calculate the original binding free energy difference. Since various errors may exist during the simulation process, loop closure correction is required to obtain the correct binding free energy difference. The original calculation results can be corrected using the loop closure correction method to eliminate errors and obtain more accurate binding free energy. For example, error analysis is performed through multiple independent simulations (repeated calculations 5 times), and the original binding free energy difference is corrected in combination with experimental data (binding free energy measured by isothermal titration calorimetry ITC) to generate a corrected energy difference. Finally, covalent molecular information is generated based on the ligand binding free energy, including detailed information such as the binding mode of the ligand and the interaction strength between the ligand and the protein.

[0048] In an embodiment of the present invention, by streamlining the molecular dynamics simulation process, the time and workload of manual operation can be greatly reduced compared to traditional manual operation, and a large number of molecular dynamics simulation tasks can be quickly completed. Through standardized processing steps and precise parameter settings, human errors are avoided, and the accuracy and reliability of the simulation results are improved. Fixed operating procedures and parameter settings ensure that each simulation has a high degree of consistency and repeatability, which facilitates data comparison and verification between researchers. At the same time, it can not only process non-covalent ligands, but also accurately simulate and calculate the binding free energy of covalent ligands, providing more powerful technical support for fields such as drug development and material design, and helping to accelerate the research process and innovative development in related fields.

[0049] See Figure 2Based on the same inventive concept, an embodiment of the present invention provides a covalent molecular simulation device based on AMBER, which is applied to the method provided by an embodiment of the present invention. The device includes: a structure acquisition unit for acquiring a protein structure; a system construction unit for constructing an initialization simulation system based on AMBER and the protein structure; a simulation operation unit for performing a simulation operation based on the initialization simulation system to obtain simulation operation data; a configuration unit for topologically configuring the simulation operation data to generate configured data; and an information generation unit for generating covalent molecular information based on the configured data.

[0050] Furthermore, an embodiment of the present invention also provides a computer-readable storage medium having a computer program stored thereon, which implements the method described in the embodiment of the present invention when the program is executed by a processor.

[0051] The above describes in detail the optional implementation methods of the embodiments of the present invention in conjunction with the accompanying drawings. However, the embodiments of the present invention are not limited to the specific details in the above implementation methods. Within the technical concept of the embodiments of the present invention, various simple modifications can be made to the technical solutions of the embodiments of the present invention, and these simple modifications all fall within the scope of protection of the embodiments of the present invention.

[0052] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any appropriate manner without contradiction. To avoid unnecessary repetition, the embodiments of the present invention will not further describe various possible combinations.

[0053] Those skilled in the art will understand that all or part of the steps in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a program, which is stored in a storage medium and includes a number of instructions for causing a single-chip microcomputer, chip or processor to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program code.

[0054] In addition, various implementations of the embodiments of the present invention may be arbitrarily combined, and as long as they do not violate the concept of the embodiments of the present invention, they should also be regarded as the contents disclosed in the embodiments of the present invention.

Claims

1. A covalent molecular simulation method based on AMBER, characterized in that: The method comprises: Obtain protein structure; Constructing an initial simulation system based on AMBER and the protein structure; Executing simulation operation based on the initialized simulation system to obtain simulation operation data; Performing topological configuration on the simulation operation data to generate configured data; Covalent molecule information is generated based on the configured data.

2. The method according to claim 1, characterized in that The initialization simulation system based on AMBER and the protein structure is constructed, comprising: Preprocessing the protein structure to obtain a preprocessed structure; Inputting the preprocessed structure into AMBER, loading the corresponding force field data, performing charge neutralization processing, and obtaining a first initial system; Performing a solvation process on the first initial system based on a preset water molecule model to obtain a processed system; The treated system is subjected to ion concentration configuration to obtain an initialized simulation system.

3. The method according to claim 2, characterized in that The preprocessing of the protein structure to obtain the preprocessed structure comprises: performing redundancy analysis on the protein structure to obtain redundant chains / heteroatoms; Processing the protein structure based on the redundant chains / heteroatoms and a preset crystal water threshold to obtain a first processed structure; obtaining a homology model, and performing protonation on the first processed structure based on the homology model to obtain a second processed structure; performing residue atom analysis on the second processed structure to obtain unidentified residue atoms; The second processed structure is processed based on the unidentified residue atoms to generate a preprocessed structure.

4. The method according to claim 1, wherein The performing simulation operation based on the initialized simulation system to obtain simulation operation data includes: performing energy minimization processing on the initialized simulation system to obtain a process or system; performing a simulated heating process based on the treated system to obtain a heated system; performing a balancing process on the heated system to obtain a balanced system; A production simulation is performed based on the balanced system to obtain simulation operation data.

5. The method according to claim 1, wherein The performing topological configuration on the simulation operation data to generate configured data includes: extracting non-covalent ligands and covalent ligands from the simulation run data; performing parameterization processing on the non-covalent ligand to obtain first processed data; performing parameterization processing on the covalent ligand to obtain second processed data; Configured data is generated based on the first processed data and the second processed data.

6. The method according to claim 5, characterized in that The performing parameterization processing on the non-covalent ligand to obtain first processed data includes: Obtain all first non-ligand data in the non-covalent ligand; Deleting the first non-ligand data from the non-covalent ligand to obtain first post-deletion data; performing proton ligand perfection processing on the first deleted data to obtain perfected data; Normative processing is performed on the improved data to generate first processed data.

7. The method according to claim 5, characterized in that The performing parameterization processing on the covalent ligand to obtain second processed data includes: obtaining second non-ligand data from the covalent ligand based on the covalent bond; Deleting the second non-ligand data from the covalent ligand to obtain second deleted data; performing a protonation process on the second deleted data to obtain protonated data; performing a ligation chain process on the protonated data to obtain ligated data; performing charge distribution analysis on the connected data to obtain charge data; Second processed data is generated based on the connected data and the charge data.

8. The method according to claim 1, characterized in that The generating of covalent molecule information based on the configured data includes: performing maximum common substructure analysis on the configured data to generate analysis data; Performing atomic transformation analysis on the analysis data based on energy conservation rules to generate closed loop information; Calculating and generating an original binding free energy difference based on the closed loop information; Correcting the original binding free energy difference to generate a corrected energy difference; determining a ligand binding free energy based on the corrected energy difference; Covalent molecular information is generated based on the ligand binding free energy.

9. A covalent molecular simulation device based on AMBER, characterized in that The device is applied to the method according to any one of claims 1 to 8, and the device comprises: Structure acquisition unit, used to obtain protein structure; A system construction unit, used for constructing an initialization simulation system based on AMBER and the protein structure; A simulation operation unit, configured to execute a simulation operation based on the initialized simulation system and obtain simulation operation data; A configuration unit, configured to perform topological configuration on the simulation operation data to generate configured data; An information generating unit is used to generate covalent molecule information based on the configured data.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.