A method, device, medium and program product for calculating the binding free energy
Through the method of data clustering of molecular dynamics simulation trajectory files, calculation combined with free energy, the problems of computing resource limitation and slow calculation speed in the prior art are solved, and efficient and accurate calculation results are achieved.
Patent Information
- Application Number
- CN202411328683.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-23
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-09-23
AI Technical Summary
When computing and free energy are combined with the prior art, due to the computing resources and system size, the calculation accuracy is not high and the calculation speed is slow.
Through the trajectory file based on molecular dynamics simulation, the data information of each frame, including distance distribution information, is determined, and data clustering is performed, and the combined free energy is calculated based on the clustering results.
While ensuring computing accuracy, it saves a lot of computing resources and improves computing efficiency.
Smart Images

Figure CN119252357B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of bioinformatics, and particularly to a technology for calculating binding free energy. Background Art
[0002] Accurately calculating the binding free energy is an important issue in the field of molecular simulation. The free energy change between the initial state (reactants) and the final state (products) determines the extent of biological processes. Classical methods for calculating binding free energy include Free Energy Perturbation (FEP) and Thermodynamic Integration (TI). These two methods are mature methods with relatively accurate calculation results, but they require a large amount of data collection, and there are limitations in computing resources and system size in practical applications. Molecular Mechanics - Poisson Boltzmann Surface Area (MM - PBSA) is a method for estimating binding free energy by post - processing the trajectory of molecular dynamics simulation. Although the accuracy of the MM - PBSA method is not as good as that of FEP and TI, this method has a small computational amount and is an effective method in molecular recognition and differentiating the strength of binding, and has been successfully used in many calculations of binding free energy. Currently, when using the MM - PBSA method to calculate binding free energy, generally some special frames are manually selected for calculation, or all frames or some frames extracted at set intervals are calculated. The former is affected by humans and the selected quantity is limited, resulting in low calculation accuracy; while the latter has a large computational amount, requires computing resources, and has a slow calculation speed. Summary of the Invention
[0003] An object of this application is to provide a method, device, medium, and program product for calculating binding free energy.
[0004] According to one aspect of this application, a method for calculating binding free energy is provided. The method includes:
[0005] Based on the trajectory file of molecular dynamics simulation, determining the data information corresponding to each frame in the trajectory file, where the data information corresponding to each frame in the trajectory file includes distance distribution information;
[0006] Based on the data information corresponding to each frame in the trajectory file, determining one or more data clusters;
[0007] Based on the one or more data clusters, determining the corresponding binding free energy.
[0008] According to one aspect of the present application, there is provided a computer device for calculating binding free energy, including a memory, a processor, and a computer program stored on the memory, wherein the processor executes the computer program to implement the steps of any one of the above methods.
[0009] According to one aspect of the present application, there is provided a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of any one of the above methods.
[0010] According to one aspect of the present application, there is provided a computer program product including a computer program, wherein the computer program, when executed by a processor, implements the steps of any one of the above methods.
[0011] According to one aspect of the present application, there is provided a device for calculating binding free energy, the device including:
[0012] a first module configured to determine data information corresponding to each frame in the trajectory file based on the trajectory file of the molecular dynamics simulation, and the data information corresponding to each frame in the trajectory file includes distance distribution information;
[0013] a second module configured to determine one or more data clusters based on the data information corresponding to each frame in the trajectory file;
[0014] a third module configured to determine the corresponding binding free energy based on the one or more data clusters.
[0015] Compared with the prior art, the present application determines the data information corresponding to each frame in the trajectory file based on the trajectory file of the molecular dynamics simulation, and the data information corresponding to each frame in the trajectory file includes distance distribution information; determines one or more data clusters based on the data information corresponding to each frame in the trajectory file; and determines the corresponding binding free energy based on the one or more data clusters. By clustering each frame of the trajectory file and calculating the binding free energy based on the clustering result, a large amount of computing resources are saved and the computing efficiency is improved while ensuring the computing accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects, and advantages of the present application will become more apparent:
[0017] Figure 1 A flowchart showing a method for calculating binding free energy according to an embodiment of the present application;
[0018] Figure 2Shows a flowchart of a method for clustering each frame of a trajectory file according to an embodiment of the present application;
[0019] Figure 3 Shows a flowchart of a method for calculating the binding free energy according to an embodiment of the present application;
[0020] Figure 4 Shows a structural diagram of a device for calculating the binding free energy according to an embodiment of the present application;
[0021] Figure 5 Shows an exemplary system that can be used to implement the various embodiments described in the present application.
[0022] The same or similar reference numerals in the drawings represent the same or similar components. Detailed Description of the Embodiments
[0023] The present application will be further described in detail below with reference to the accompanying drawings.
[0024] In a typical configuration of the present application, the terminal, the devices of the service network, and the trusted party each include one or more processors (e.g., a Central Processing Unit (CPU)), an input / output interface, a network interface, and a memory.
[0025] The memory may include non-permanent memory in the computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash memory. The memory is an example of the computer-readable medium.
[0026] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PCM), programmable random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device.
[0027] The devices referred to in this application include, but are not limited to, user devices, network devices, or devices formed by integrating user devices and network devices through a network. The user devices include, but are not limited to, any mobile electronic product that can perform human-computer interaction with users (such as through a touchpad), such as smart phones, tablet computers, etc. The mobile electronic products can adopt any operating system, such as the Android operating system, the iOS operating system, etc. Among them, the network device includes an electronic device that can automatically perform numerical calculations and information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, a microprocessor, an application specific integrated circuit (ASIC), a programmable logic device (PLD), a field programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc. The network device includes, but is not limited to, a computer, a network host, a single network server, a set of multiple network servers, or a cloud composed of multiple servers; here, the cloud is composed of a large number of computers or network servers based on cloud computing. Among them, cloud computing is a type of distributed computing, which consists of a virtual supercomputer composed of a group of loosely coupled computers. The network includes, but is not limited to, the Internet, a wide area network, a metropolitan area network, a local area network, a VPN network, a wireless ad hoc network (Ad Hoc network), etc. Preferably, the device can also be a program running on the user device, the network device, or the device formed by integrating the user device and the network device, the network device, the touch terminal, or the device formed by integrating the network device and the touch terminal through a network.
[0028] Of course, those skilled in the art should understand that the above devices are only examples. Other existing or future devices that may be applicable to this application should also be included within the protection scope of this application and are hereby incorporated by reference.
[0029] In the description of this application, "a plurality of" means two or more, unless otherwise specifically defined.
[0030] Figure 1A flowchart of a method for calculating the binding free energy according to an embodiment of the present application is shown. The method includes: step S11, step S12, and step S13. In step S11, device 1 determines data information corresponding to each frame in the trajectory file based on the trajectory file of molecular dynamics simulation. The data information corresponding to each frame in the trajectory file includes distance distribution information. In step S12, device 1 determines one or more data clusters based on the data information corresponding to each frame in the trajectory file. In step S13, device 1 determines the corresponding binding free energy based on the one or more data clusters.
[0031] In step S11, device 1 determines data information corresponding to each frame in the trajectory file based on the trajectory file of molecular dynamics simulation. The data information corresponding to each frame in the trajectory file includes distance distribution information.
[0032] In some embodiments, device 1 includes, but is not limited to, user devices and network devices with information processing or computing capabilities, such as, for example, tablet computers, computers, servers, etc.
[0033] In some embodiments, the molecular dynamics simulation (Molecular Dynamics Simulation) refers to using the basic principles of Newtonian mechanics to generate continuous configurations of a protein molecular system by integrating the equations of motion, thereby obtaining the movement trajectories of all atoms, that is, the process in which the positions and velocities of the atoms in the system change over time. Generally, simulation tools such as GROMACS (www.gromacs.org), AMBER (www.ambermd.org), or CHARMM (www.academiccharmm.org) can be used to perform this molecular dynamics simulation to obtain the corresponding trajectory file. The trajectory file contains detailed information about all atoms during the entire simulation process. The trajectory file is usually distinguished by frames, and each frame represents a snapshot at a specific time step. Each frame of the trajectory file usually contains time step information associated with this frame, coordinate information and velocity information of each atom at this time step, and other possible data (such as temperature, pressure, dimensions and shape of the simulation box, etc.).
[0034] In some embodiments, device 1 may calculate the data information corresponding to each frame in the trajectory file based on the trajectory file for subsequent clustering of each frame of the trajectory file. The data information corresponding to each frame in the trajectory file includes distance distribution information. The distance distribution information includes the target atom pair distance information in this frame. The target atom pair distance information may be the atom pair distance information calculated based on specific atom pairs selected according to actual needs. For example, the atom pair distance between substrate (a small molecule that interacts with a protein biomolecule in a simulation) atoms and protein atoms, the atom pair distance between a part of substrate atoms and protein atoms, the atom pair distance between substrate atoms and a part of protein atoms, the atom pair distance between a part of substrate atoms and a part of protein atoms, etc.
[0035] In some embodiments, the data information corresponding to each frame in the trajectory file further includes the root mean square deviation of the protein structure and the solvent accessible surface area. The root mean square deviation of the protein structure refers to the root mean square deviation (RMSD, Root Mean Square Deviation) of the protein structure in this frame calculated relative to the initial structure or the average structure. The solvent accessible surface area (SASA, Solvent Accessible Surface Area) refers to the surface area of the protein molecule that the solvent can contact. Here, those skilled in the art should understand that the above data information is only an example, and other existing or future possible data information that can be applied to this application should also be included within the protection scope of this application and is hereby incorporated by reference.
[0036] In some embodiments, the distance distribution information includes target atom pair distance information; step S11 includes: for each frame in the trajectory file, device 1 determines the target atoms in the substrate and the target atoms in the protein, where the target atoms in the substrate and the target atoms in the protein meet the preset screening conditions; and determines the target atom pair distance information corresponding to the target atoms in the substrate and the target atoms in the protein.
[0037] In some embodiments, corresponding screening conditions can be designed based on actual needs to select corresponding target atoms. For example, substrate target atoms and protein target atoms directly related to protein function can be selected, or substrate target atoms and protein target atoms can be selected according to different types of interactions such as hydrogen bonds, hydrophobic interactions, ionic bonds, and van der Waals forces, or substrate target atoms and protein target atoms that are close to each other in three-dimensional space can be selected. In some embodiments, the preset screening conditions include at least any one of the following: the spatial distance between the target atoms in the substrate and the target atoms in the protein satisfies a preset distance condition; the charge numbers of the target atoms in the substrate and the target atoms in the protein satisfy a preset charge number condition. For example, target atoms in the substrate and target atoms in the protein with a spatial distance less than a distance threshold (for example, the distance threshold can be set to ) are selected, and / or the charge numbers of the target atoms in the selected substrate and the target atoms in the protein are integers. After the target atoms in the substrate and the target atoms in the protein are selected, based on their coordinate information, the distance between the two is directly calculated to obtain the distance information of the target atom pair.
[0038] In step S12, the device 1 determines one or more data clusters based on the data information corresponding to each frame in the trajectory file. For example, the device 1 can cluster each frame in the trajectory file based on the data information corresponding to each frame in the trajectory file using a corresponding clustering algorithm (such as the K-means clustering algorithm, the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) clustering algorithm, the hierarchical clustering algorithm, etc.) to obtain one or more data clusters. Each data cluster includes at least one frame in the trajectory file.
[0039] In some embodiments, refer to Figure 2The flowchart of the method for clustering each frame of the trajectory file is shown. The step S12 includes: step S121, device 1 performs dimensionality reduction processing on the data information corresponding to each frame in the trajectory file to obtain the dimensionality-reduced data information; step S122, device 1 performs clustering processing on the dimensionality-reduced data information to determine one or more data clusters. In some embodiments, in order to save computing resources and improve computing efficiency, device 1 may first perform dimensionality reduction processing on the data information corresponding to each frame in the trajectory file, and then perform clustering on each frame of the trajectory file based on the dimensionality-reduced data information. In some embodiments, dimensionality reduction algorithms such as t-distributed Stochastic Neighbor Embedding (t-SNE), Uniform Manifold Approximation and Projection (UMAP), or Multiple Dimensional Scaling (MDS) may be used to perform dimensionality reduction of the data information.
[0040] In some embodiments, the step S121 includes: step S1211 (not shown), device 1 determines the conditional probability distribution of the data information corresponding to each frame in the trajectory file in the original high-dimensional space; step S1212 (not shown), device 1 determines the initialization information of the data information corresponding to each frame in the trajectory file in the low-dimensional space and the corresponding conditional probability distribution in the low-dimensional space; step S1213 (not shown), device 1 iteratively optimizes the initialization information based on the conditional probability distribution in the original high-dimensional space and the conditional probability distribution in the low-dimensional space to obtain the corresponding dimensionality-reduced data information.
[0041] For example, the device 1 can calculate the similarity between the data information corresponding to each frame in the trajectory file using methods such as Euclidean distance or cosine similarity, thereby obtaining a similarity matrix of these data information in the original high-dimensional space. The similarity matrix can be normalized and converted into a probability form to form a conditional probability distribution of the data information in the original high-dimensional space. Before iterative optimization, initialization dimensionality reduction can be performed to obtain the initial value (i.e., initialization information) corresponding to the data information in the low-dimensional space. In some embodiments, the initialization information of the data information corresponding to each frame in the trajectory file in the low-dimensional space includes: linear dimensionality reduction of the data information corresponding to each frame in the trajectory file to obtain the initialization information of the data information corresponding to each frame in the trajectory file in the low-dimensional space. For example, a linear dimensionality reduction method such as principal component analysis (PCA, Principal Components Analysis) can be used to reduce the dimensionality of the data information in the original high-dimensional space to obtain the initialization information corresponding to these data information in the low-dimensional space. In some embodiments, these data information can also be randomly initialized to obtain random initialization information in the low-dimensional space. For the initialization information obtained in the low-dimensional space, the same method as the conditional probability distribution of the data information in the original high-dimensional space can be used to calculate the corresponding conditional probability distribution in the low-dimensional space. The difference between the conditional probability distribution of the original high-dimensional space and the low-dimensional space is measured by the corresponding objective function, and the initialization information is optimized by the gradient descent algorithm. The objective function is minimized through continuous iterative optimization, and finally the data information after dimensionality reduction is obtained.
[0042] In some embodiments, step S1213 includes: device 1 determines the relative entropy between the conditional probability distribution in the original high-dimensional space and the conditional probability distribution in the low-dimensional space; iteratively optimizes the initialization information by minimizing the relative entropy to obtain the corresponding data information after dimensionality reduction. For example, the difference between the original high-dimensional space and the conditional probability distribution in the low-dimensional space can be measured by relative entropy. Use the gradient descent algorithm to optimize the data information corresponding to the information in the low-dimensional space to minimize the relative entropy. Continue to iterate and optimize until the corresponding stop condition is met (for example, the maximum number of iterations is reached or the relative entropy change between two consecutive iterations is less than a preset threshold, etc.), and obtain the final data information after dimensionality reduction.
[0043] In step S13, the device 1 determines the corresponding binding free energy based on the one or more data clusters. For example, each data cluster includes at least one frame in the trajectory file. The device 1 can select a frame from each data cluster. The binding free energy corresponding to the selected frame in each data cluster is calculated by the MM-PBSA method. Based on the binding free energy corresponding to the selected frame in each data cluster, the final binding free energy is determined.
[0044] In some embodiments, referring to Figure 3 the flowchart of the method for calculating the binding free energy shown, step S13 includes: step S131 (not shown), the device 1 determines the target frame information corresponding to each data cluster in the one or more data clusters; step S132 (not shown), the device 1 determines the corresponding binding free energy based on the target frame information. In some embodiments, the target frame information includes the target frame corresponding to each data cluster. The target frame can be a frame randomly selected from the data cluster; or it can be the frame closest to the cluster center in the data cluster. The device 1 determines the binding free energy corresponding to each target frame information through the MM-PBSA method, and then determines the corresponding binding free energy. For example, the mean value of the binding free energies corresponding to each target frame information can be taken; or the weights corresponding to each target frame are determined, and the binding free energy is determined in combination with the weights.
[0045] In some embodiments, the target frame information further includes the weight information corresponding to the target frame; step S132 includes: determining the binding free energy corresponding to the target frame; determining the corresponding binding free energy based on the binding free energy corresponding to the target frame and the weight information corresponding to the target frame. In some embodiments, the ratio of the number of frames included in the data cluster corresponding to the target frame to the total number of frames can be used as the weight information corresponding to the target frame. The inner product of the binding free energy of the target frame determined by the MM-PBSA method and the corresponding weight information is taken to obtain the final binding free energy of the trajectory.
[0046] In addition, the inventors conducted a set of test experiments with 6 trajectory files, each with 250 frames. Table 1 is a comparison table of the calculation results of the method described in the present application and the traditional calculation of all frames. Table 2 is a comparison table of the calculation time consumption of the two methods. It can be seen from the following two tables that while ensuring the calculation accuracy, the present application can greatly improve the calculation efficiency.
[0047] Table 1 Comparison Table of Calculation Results
[0048] Trajectory file The method of this application Calculate all frames Calculation result difference D09 -14.04 -14.09 0.05 D10 -16.54 -16.24 -0.30 D11 -16.86 -17.01 0.15 D13 -19.13 -18.55 -0.58 D14 -21.51 -21.54 0.03 D15 -23.61 -23.89 0.28
[0049] Table 2 Comparison Table of Calculation Time Consumption
[0050]
[0051] Figure 4The structural diagram of a device for calculating the binding free energy according to an embodiment of the present application is shown. The device 1 includes a module 11, a module 12, and a module 13. The module 11 determines the data information corresponding to each frame in the trajectory file based on the trajectory file of the molecular dynamics simulation. The data information corresponding to each frame in the trajectory file includes distance distribution information. The module 12 determines one or more data clusters based on the data information corresponding to each frame in the trajectory file. The module 13 determines the corresponding binding free energy based on the one or more data clusters. Herein, the Figure 4 The specific implementation manners corresponding to the shown module 11, module 12, and module 13 are the same as or similar to the specific embodiments of the foregoing steps S11, S12, and S13 respectively, and thus will not be described in detail and are included herein by reference.
[0052] In some embodiments, the module 12 includes a unit 121 and a unit 122. The unit 121 performs dimensionality reduction processing on the data information corresponding to each frame in the trajectory file to obtain the dimensionality-reduced data information. The unit 122 performs clustering processing on the dimensionality-reduced data information to determine one or more data clusters. Herein, the specific implementation manners of the unit 121 and the unit 122 are the same as or similar to the specific embodiments of the foregoing steps S121 and S122 respectively, and thus will not be described in detail and are included herein by reference.
[0053] In some embodiments, the unit 121 includes a subunit 1211 (not shown), a subunit 1212 (not shown), and a subunit 1213 (not shown). The subunit 1211 determines the conditional probability distribution of the data information corresponding to each frame in the trajectory file in the original high-dimensional space. The subunit 1212 determines the initialization information of the data information corresponding to each frame in the trajectory file in the low-dimensional space and the corresponding conditional probability distribution in the low-dimensional space. The subunit 1213 iteratively optimizes the initialization information based on the conditional probability distribution in the original high-dimensional space and the conditional probability distribution in the low-dimensional space to obtain the corresponding dimensionality-reduced data information. Herein, the specific implementation manners of the subunit 1211, the subunit 1212, and the subunit 1213 are the same as or similar to the specific embodiments of the foregoing steps S1211, S1212, and S1213 respectively, and thus will not be described in detail and are included herein by reference.
[0054] In some embodiments, the 1-3 module 13 includes a 1-3-1 unit 131 and a 1-3-2 unit 132. The 1-3-1 unit 131 determines the target frame information corresponding to each data cluster in the one or more data clusters; the 1-3-2 unit 132 determines the corresponding binding free energy based on the target frame information. Here, the specific implementation manners of the 1-3-1 unit 131 and the 1-3-2 unit 132 are the same as or similar to the specific embodiments of the foregoing step S131 and step S132 respectively, so they will not be described in detail and are included herein by reference.
[0055] Figure 5 An exemplary system that can be used to implement the various embodiments described in the present application is shown.
[0056] As Figure 5 shown, in some embodiments, the system 300 can act as any one of the devices in the various embodiments. In some embodiments, the system 300 may include one or more computer-readable media having instructions (e.g., system memory or NVM / storage device 320) and one or more processors (e.g., (one or more) processors 305) coupled to the one or more computer-readable media and configured to execute the instructions to implement modules to perform the actions described in the present application.
[0057] For one embodiment, the system control module 310 may include any suitable interface controller to provide any suitable interface to at least one of the (one or more) processors 305 and / or any suitable device or component communicating with the system control module 310.
[0058] The system control module 310 may include a memory controller module 330 to provide an interface to the system memory 315. The memory controller module 330 may be a hardware module, a software module, and / or a firmware module.
[0059] The system memory 315 may be used, for example, to load and store data and / or instructions for the system 300. For one embodiment, the system memory 315 may include any suitable volatile memory, e.g., suitable DRAM. In some embodiments, the system memory 315 may include double data rate type four synchronous dynamic random access memory (DDR4 SDRAM).
[0060] For one embodiment, the system control module 310 may include one or more input / output (I / O) controllers to provide an interface to the NVM / storage device 320 and the (one or more) communication interfaces 325.
[0061] For example, the NVM / storage device 320 can be used to store data and / or instructions. The NVM / storage device 320 can include any suitable non-volatile memory (e.g., flash memory) and / or can include any suitable (one or more) non-volatile storage devices (e.g., one or more hard disk drives (HDDs), one or more compact disc (CD) drives, and / or one or more digital versatile disc (DVD) drives).
[0062] The NVM / storage device 320 can include storage resources that are physically part of the device on which the system 300 is installed, or it can be accessed by the device without being part of the device. For example, the NVM / storage device 320 can be accessed via a network through the (one or more) communication interfaces 325.
[0063] (One or more) communication interfaces 325 can provide an interface for the system 300 to communicate through one or more networks and / or with any other suitable device. The system 300 can wirelessly communicate with one or more components of a wireless network according to any of one or more wireless network standards and / or protocols.
[0064] For one embodiment, at least one of the (one or more) processors 305 can be logically encapsulated with one or more controllers of the system control module 310 (e.g., the memory controller module 330). For one embodiment, at least one of the (one or more) processors 305 can be logically encapsulated with one or more controllers of the system control module 310 to form a system-in-package (SiP). For one embodiment, at least one of the (one or more) processors 305 can be logically integrated with one or more controllers of the system control module 310 on the same die. For one embodiment, at least one of the (one or more) processors 305 can be logically integrated with one or more controllers of the system control module 310 on the same die to form a system-on-chip (SoC).
[0065] In various embodiments, the system 300 can be, but is not limited to: a server, a workstation, a desktop computing device, or a mobile computing device (e.g., a laptop computing device, a handheld computing device, a tablet computer, a netbook, etc.). In various embodiments, the system 300 can have more or fewer components and / or a different architecture. For example, in some embodiments, the system 300 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touchscreen display), a non-volatile memory port, multiple antennas, a graphics chip, an application specific integrated circuit (ASIC), and speakers.
[0066] In addition to the methods and devices described in the above embodiments, the present application also provides a computer-readable storage medium storing computer code, and when the computer code is executed, the method described in any of the previous items is executed.
[0067] The present application also provides a computer program product, and when the computer program product is executed by a computer device, the method described in any of the previous items is executed.
[0068] The present application also provides a computer device, which includes:
[0069] One or more processors;
[0070] A memory for storing one or more computer programs;
[0071] When the one or more computer programs are executed by the one or more processors, the one or more processors are caused to implement the method described in any of the previous items.
[0072] It should be noted that the present application can be implemented in software and / or a combination of software and hardware. For example, it can be implemented using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In one embodiment, the software program of the present application can be executed by a processor to implement the steps or functions described above. Similarly, the software program (including related data structures) of the present application can be stored in a computer-readable recording medium, such as a RAM memory, a magnetic or optical drive, or a floppy disk and similar devices. Additionally, some steps or functions of the present application can be implemented using hardware, for example, as a circuit that cooperates with a processor to execute each step or function.
[0073] In addition, a part of the present application can be applied as a computer program product, such as computer program instructions, and when executed by a computer, through the operation of the computer, it can call or provide the methods and / or technical solutions according to the present application. Those skilled in the art should be able to understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to the computer.
[0074] A communication medium includes a medium through which communication signals that include, for example, computer-readable instructions, data structures, program modules, or other data are transmitted from one system to another. The communication medium can include wired transmission media (such as cables and wires (e.g., optical fibers, coaxial, etc.)) and wireless (unguided) media that can propagate energy waves, such as acoustic, electromagnetic, RF, microwave, and infrared. The computer-readable instructions, data structures, program modules, or other data can be embodied as, for example, a modulated data signal in a wireless medium (such as a carrier wave or a similar mechanism that is part of what is embodied as spread spectrum technology). The term "modulated data signal" refers to a signal in which one or more of its characteristics are changed or set in a manner that encodes information in the signal. Modulation can be an analog, digital, or hybrid modulation technique.
[0075] By way of example, and not limitation, computer-readable storage media can include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules, or other data. For example, computer-readable storage media includes, but is not limited to, volatile memory such as random access memory (RAM, DRAM, SRAM); and non-volatile memory such as flash memory, various read-only memories (ROM, PROM, EPROM, EEPROM), magnetic and ferromagnetic / ferroelectric memories (MRAM, FeRAM); and magnetic and optical storage devices (hard disks, tapes, CDs, DVDs); or other media now known or later developed that are capable of storing computer-readable information / data for use by a computer system.
[0076] Here, an embodiment according to the present application includes an apparatus that includes a memory for storing computer program instructions and a processor for executing the program instructions, wherein, when the computer program instructions are executed by the processor, the apparatus is triggered to operate based on the methods and / or technical solutions according to the foregoing multiple embodiments of the present application.
[0077] For those skilled in the art, it is obvious that the present application is not limited to the details of the above-mentioned exemplary embodiments, and the present application can be implemented in other specific forms without departing from the spirit or basic characteristics of the present application. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present application is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be encompassed within the present application. Any reference signs in the claims should not be construed as limiting the claimed rights. In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices recited in the apparatus claims can also be implemented by one unit or device through software or hardware. The words such as "first" and "second" are used to denote names and do not denote any particular order.
Claims
1. A method for calculating binding free energy, wherein: The method comprises: Based on the trajectory file of molecular dynamics simulation, determining data information corresponding to each frame in the trajectory file, wherein the data information corresponding to each frame in the trajectory file includes distance distribution information, wherein the distance distribution information includes target atom pair distance information, and the target atom pair distance information is used to cluster each frame of the trajectory file; Determine one or more data clusters based on data information corresponding to each frame in the trajectory file; Based on the one or more data clusters, determining the corresponding binding free energy, wherein the determining the corresponding binding free energy based on the one or more data clusters includes: determining target frame information corresponding to each data cluster in the one or more data clusters; and determining the corresponding binding free energy based on the target frame information.
2. The method according to claim 1, wherein: The trajectory file based on molecular dynamics simulation determines the data information corresponding to each frame in the trajectory file, wherein the data information corresponding to each frame in the trajectory file includes distance distribution information including: For each frame in the trajectory file, determining a target atom in the substrate and a target atom in the protein, wherein the target atom in the substrate and the target atom in the protein meet a preset screening condition; Determine target atom pair distance information between the target atom in the substrate and the target atom in the protein.
3. The method according to claim 2, wherein: The preset screening condition includes at least one of the following: The spatial distance between the target atom in the substrate and the target atom in the protein satisfies a preset distance condition; The charge numbers of the target atom in the substrate and the target atom in the protein meet a preset charge number condition.
4. The method according to any one of claims 1 to 3, wherein: The determining one or more data clusters based on the data information corresponding to each frame in the trajectory file includes: Performing dimensionality reduction processing on the data information corresponding to each frame in the trajectory file to obtain the data information after dimensionality reduction; The data information after dimension reduction is clustered to determine one or more data clusters.
5. The method according to claim 4, wherein: The performing dimensionality reduction processing on the data information corresponding to each frame in the trajectory file to obtain the data information after dimensionality reduction includes: Determine the conditional probability distribution of the data information corresponding to each frame in the trajectory file in the original high-dimensional space; Determine initialization information of data information corresponding to each frame in the trajectory file in the low-dimensional space and the corresponding conditional probability distribution in the low-dimensional space; Based on the conditional probability distribution in the original high-dimensional space and the conditional probability distribution in the low-dimensional space, the initialization information is iteratively optimized to obtain corresponding data information after dimensionality reduction.
6. The method according to claim 5, wherein: The iterative optimization of the initialization information based on the conditional probability distribution in the original high-dimensional space and the conditional probability distribution in the low-dimensional space to obtain the corresponding data information after dimensionality reduction includes: Determine the relative entropy between the conditional probability distribution in the original high-dimensional space and the conditional probability distribution in the low-dimensional space; The initialization information is iteratively optimized by minimizing the relative entropy to obtain corresponding data information after dimensionality reduction.
7. The method according to claim 5 or 6, wherein: The determining of the initialization information of the data information corresponding to each frame in the trajectory file in the low-dimensional space includes: Linear dimensionality reduction is performed on the data information corresponding to each frame in the trajectory file to obtain initialization information of the data information corresponding to each frame in the trajectory file in a low-dimensional space.
8. The method according to claim 1, wherein: The target frame information includes a target frame corresponding to each data cluster, and the target frame is closest to a cluster center in the data cluster.
9. The method according to claim 8, wherein: The target frame information also includes weight information corresponding to the target frame; Determining the corresponding binding free energy based on the target frame information includes: determining a binding free energy corresponding to the target frame; Based on the binding free energy corresponding to the target frame and the weight information corresponding to the target frame, the corresponding binding free energy is determined.
10. The method according to claim 1, wherein: The data information corresponding to each frame in the trajectory file also includes the root mean square deviation of the protein structure and the solvent accessible surface area.
11. A computer device for calculating binding free energy, comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 10.
12. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
13. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.