Method, device, medium and program product for determining acidity coefficient

By constructing a simulation box containing proteins, solvents, and buffers, and using machine learning force fields for molecular dynamics simulation, the problems of insufficient accuracy and high cost in calculating protein acidity coefficients in existing technologies have been solved, achieving efficient and accurate acidity coefficient calculation.

CN121393584APending Publication Date: 2026-01-23SHANGHAI MOLECULAR HEART INTELLIGENT TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511613971.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-05
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

When calculating the acidity coefficients of dissociable amino acid groups in proteins, the classical force field method is not accurate enough, while the ab initio method is too costly and difficult to achieve efficient simulation in complex systems.

Method used

A simulation box containing proteins, solvents, and buffers was constructed. Molecular dynamics simulations were performed using a trained machine learning force field to obtain trajectory file information, which was then used to determine the acidity coefficient.

Benefits of technology

This technology enables high-precision and low-cost calculation of acidity coefficients in protein solution systems, improving the efficiency and accuracy of calculations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121393584A_ABST
    Figure CN121393584A_ABST
Patent Text Reader

Abstract

The invention aims to provide a method, equipment, medium and program product for determining acidity coefficient, the method comprises the following steps: constructing a simulation box corresponding to protein, the simulation box comprising the protein, a corresponding solvent and a buffer solution; performing molecular dynamics simulation on the simulation box by utilizing a trained machine learning force field to obtain corresponding track file information; and determining an acidity coefficient corresponding to each drippable site in the protein based on the track file information. A dynamic process of protein in a solution is accurately and efficiently simulated by utilizing a machine learning force field, and an accurate acidity coefficient is calculated based on a state of a protein solution system in simulation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of bioinformatics, and in particular to a technology for determining acidity coefficient. BACKGROUND

[0002] Acidity coefficient (pKa) is used to measure the dissociation tendency of each group of amino acids. By comparing the pH value of the environment with the pKa of the corresponding amino acid, the (de) protonation state of the amino acid can be predicted. For amino acids in proteins, due to the influence of the (de) protonation state of other amino acids and the structure of the protein, the corresponding pKa of the amino acid will often change compared to the pKa of the amino acid in an isolated solution environment. Determining the pKa of each group of dissociable (titratable) amino acids in a protein can help understand the structure and function of the protein and guide protein design.

[0003] The calculation of the acidity coefficient (pKa) of each group of dissociable (titratable) amino acids in a protein mainly relies on the classical force field (Classical Force Fields) method and the ab initio calculation method based on quantum chemical principles. The classical force field method uses empirical or semi-empirical function forms to approximately describe the potential energy surface, thereby realizing efficient simulation of biological macromolecules, material systems, and condensed phases. However, the empirical parameters often ignore or incorrectly represent key physical effects such as polarization effect, charge transfer, and electron cloud rearrangement, affecting the calculation accuracy. The ab initio calculation method directly solves the electronic Schrödinger equation, which can construct the potential energy surface (PES) with high accuracy. However, this method has high computational cost, and when dealing with large systems or long time scales, the computational power will quickly exceed the range of tolerance, which makes it lack of practicality in common simulation of complex systems such as protein solution systems. SUMMARY

[0004] An object of the present application is to provide a method, device, medium and program product for determining acidity coefficient.

[0005] According to one aspect of the present application, a method for determining acidity coefficient is provided, the method comprising:

[0006] constructing a simulation box corresponding to the protein, wherein the simulation box comprises the protein, a corresponding solvent and a buffer;

[0007] performing molecular dynamics simulation on the simulation box using the trained machine learning force field to obtain corresponding trajectory file information;

[0008] determining the acidity coefficient corresponding to each titratable site in the protein based on the trajectory file information.

[0009] According to an aspect of the present application, there is provided a computer device for determining acidity coefficient, comprising a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of any of the methods described above.

[0010] According to an aspect of the present application, there is provided a computer readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of any of the methods described above.

[0011] According to an aspect of the present application, there is provided a computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the methods described above.

[0012] According to an aspect of the present application, there is provided a device for determining acidity coefficient, comprising:

[0013] a first module for constructing a simulation box corresponding to a protein, wherein the simulation box comprises the protein, a corresponding solvent and a buffer;

[0014] a second module for performing molecular dynamics simulation on the simulation box by using a trained machine learning force field to obtain corresponding trajectory file information;

[0015] a third module for determining an acidity coefficient corresponding to each titratable site in the protein based on the trajectory file information.

[0016] Compared with the prior art, the present application constructs a simulation box corresponding to a protein, wherein the simulation box comprises the protein, a corresponding solvent and a buffer; performs molecular dynamics simulation on the simulation box by using a trained machine learning force field to obtain corresponding trajectory file information; and determines an acidity coefficient corresponding to each titratable site in the protein based on the trajectory file information. The machine learning force field is used to accurately and efficiently simulate the dynamic process of the protein in the solution, and the accurate acidity coefficient is calculated based on the state of the protein solution system in the simulation. BRIEF DESCRIPTION OF DRAWINGS

[0017] Other features, objects and advantages of the present application will become more apparent from the following detailed description of non-limiting embodiments, made with reference to the attached drawings:

[0018] Figure 1 a flow chart of a method for determining acidity coefficient according to an embodiment of the present application is shown;

[0019] Figure 2 a structural diagram of a device for determining acidity coefficient according to an embodiment of the present application is shown;

[0020] Figure 3 An example system that can be used to implement various embodiments described herein is shown.

[0021] The same or similar reference numerals can represent the same or similar elements throughout the drawings. DETAILED DESCRIPTION

[0022] The application is described in further detail below with reference to the accompanying drawings.

[0023] In one example configuration of the application, the terminal, the device of the service network and the trusted party each include one or more processors (e.g., a Central Processing Unit (CPU)), input / output interfaces, network interfaces, and memory.

[0024] The memory can include non-persistent memory and / or volatile memory, such as a random access memory (RAM) and / or a non-volatile memory, such as a read-only memory (ROM), EPROM, EEPROM, or flash memory. The memory is an example of computer readable media.

[0025] Computer readable media includes permanent and non-permanent, movable and non-movable media that can be implemented by any method or technology for storing information. The information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PCM), programmable random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other memory technology, compact disc read only memory (Compact Disc Read-Only Memory, CD-ROM), digital versatile disc (Digital Versatile Disc, DVD) or other optical storage, magnetic cassette, magnetic tape disk storage or other magnetic storage device, or any other non-transmission medium that can be used to store information that can be accessed by a computing device.

[0026] The device referred to in the present application includes but is not limited to a user device, a network device, or a device formed by integrating the user device and the network device through a network. The user device includes but is not limited to any kind of mobile electronic product capable of human-computer interaction (for example, human-computer interaction through a touch panel), such as a smart phone, a tablet computer, etc., which can adopt any operating system, such as an Android operating system, an iOS operating system, etc. The network device includes an electronic device capable of automatically performing numerical calculation and information processing according to a pre-set or stored instruction, and its hardware includes but is not limited to a microprocessor, an Application Specific Integrated Circuit (ASIC), a Programmable Logic Device (PLD), a Field Programmable Gate Array (FPGA), a Digital Signal Processor (DSP), an embedded device, etc. The network device includes but is not limited to a computer, a network host, a single network server, a plurality of network servers, or a cloud formed by a plurality of servers; here, the cloud is formed by a large number of computers or network servers based on cloud computing, wherein the cloud computing is a kind of distributed computing, and a virtual supercomputer formed by a group of loosely coupled computer clusters. The network includes but is not limited to the Internet, a wide area network, a metropolitan area network, a local area network, a VPN network, a wireless Ad Hoc network, etc. Preferably, the device can also be a program running on the user device, the network device, or a device formed by integrating the user device and the network device, the network device, a touch terminal, or a device formed by integrating the network device and the touch terminal through a network.

[0027] Of course, those skilled in the art should understand that the above device is only an example, and other existing or future devices, such as devices applicable to the present application, should also be included in the protection scope of the present application, and are hereby included by reference.

[0028] In the description of the present application, the meaning of "a plurality of" is two or more, unless otherwise explicitly and specifically limited.

[0029] Figure 1A method for determining acidity coefficient is shown in a flow chart according to one embodiment of the present application, which comprises steps S11, S12 and S13. In step S11, a simulation box corresponding to a protein is constructed, wherein the simulation box contains the protein, a corresponding solvent and a buffer; in step S12, molecular dynamics simulation is performed on the simulation box by using a trained machine learning force field to obtain corresponding trajectory file information; and in step S13, based on the trajectory file information, acidity coefficients corresponding to each titratable site in the protein are determined.

[0030] In some embodiments, the device 1 includes but is not limited to a user device, a network device, etc. having information processing or computing capability, such as a tablet computer, a computer, a server, etc.

[0031] In step S11, a simulation box corresponding to a protein is constructed, wherein the simulation box contains the protein, a corresponding solvent and a buffer. The simulation box is a virtual periodic container for molecular dynamics simulation. The simulation box can be constructed based on the structure information of the protein by using a corresponding computational chemistry tool (for example, AmberTools (https: / / ambermd.org / ), GROMACS (https: / / www.gromacs.org / ), or PACKMOL (https: / / m3g.github.io / packmol / ), etc.). The protein, the corresponding solvent and the buffer for simulating the protein solution system are all placed in the simulation box. The solvent is selected based on the solution environment to be simulated. Generally, a water solution environment is simulated, i.e., water is selected as the solvent. Of course, a non-water solvent can also be selected for simulation based on simulation requirements. The buffer is used to maintain the pH environment in the simulation box stable. A buffer with a matching effective buffer range can be selected based on the pH environment to be simulated. The pH of the pH environment to be simulated can be set based on the theoretical acidity coefficient of the titratable site for which the acidity coefficient is to be determined. The theoretical acidity coefficient is the acidity coefficient of the amino acid corresponding to the titratable site in an isolated solution environment (i.e., a free amino acid). For example, the pKa of a free histidine (HIS) side chain is about 6, and if the corresponding pKa in the protein is to be determined, the pH of the simulated pH environment can be set to about 6.

[0032] In some embodiments, the step S11 comprises: a step S111 (not shown) of placing the protein in the simulation box and adding a corresponding solvent based on structural information of the protein by the device 1; and a step S112 (not shown) of replacing part of the solvent molecules added in the simulation box with buffer molecules based on the volume of the simulation box and a preset buffer concentration. In some embodiments, the structural information of the protein can be obtained from its crystal structure. For example, it can be obtained by querying a protein database (e.g., a PDB database (Protein Data Bank), a UniProt database, etc.). It can also be obtained by prediction using a corresponding protein structure prediction model (e.g., an AlphaFold series model, a RoseTTAFold model, or an ESMFold model, etc.). Generally, the protein can be placed in the central region of the simulation box based on its structural information. Due to the periodic boundary conditions (PBC) of the simulation box, it can also be placed in other positions in the simulation box. The distance between the surface of the protein in the simulation box and the boundary of the simulation box is greater than a corresponding preset distance, so as to leave enough space to fill the solvent. Then the solvent is added to the simulation box. Subsequently, the number of buffer molecules that need to be added to the simulation box can be calculated based on the volume of the simulation box and the preset buffer concentration, and a corresponding number of solvent molecules in the simulation box are replaced with buffer molecules. The size of the simulation box can be determined based on the coordinates of the solvent molecules after the solvent is added, and then the volume of the simulation box can be calculated.

[0033] In some embodiments, the step S112 comprises: i. the device 1 randomly selects one solvent molecule from the solvent added in the simulation box, replaces it with a buffer molecule corresponding to the buffer solution; determines whether the buffer molecule satisfies a replacement condition, if the condition is not satisfied, cancels the replacement, restores the buffer molecule to a solvent molecule, and returns to step i; otherwise, continues to perform step i until the number of buffer molecules added in the simulation box reaches a set number, wherein the set number is determined based on the volume of the simulation box and a preset buffer concentration. In some embodiments, the replacement condition comprises: the minimum distance between the buffer molecule and the protein is greater than or equal to a corresponding distance threshold; and the buffer molecule is located inside the simulation box. In some embodiments, the distance threshold can be set as the minimum distance that does not collide with the protein. If the minimum distance between the buffer molecule and the protein is less than the distance threshold, i.e., the buffer molecule collides with the protein, and / or the buffer molecule extends out of the simulation box, the replacement condition is not satisfied, the replacement fails, the replaced buffer molecule is restored to the original solvent molecule, and a solvent molecule is reselected for replacement. If the replacement condition is satisfied, the next replacement can be performed until the number of buffer molecules added in the simulation box reaches the set number.

[0034] In step S12, molecular dynamics simulation is performed on the simulation box using the trained machine learning force field to obtain corresponding trajectory file information. The machine learning force field (MLFF) is a force field that uses a machine learning model to fit the potential energy of the interaction between atoms. Compared with the classical force field, it can accurately capture more complex interactions and perform high-precision simulation. The size of the simulation system and the dynamic simulation time scale are larger than those of the ab initio method, and it is more practical. By performing molecular dynamics simulation on the simulation box using the machine learning force field, the position, velocity, force, etc. of each atom in the simulation box can be obtained over time, thereby obtaining the trajectory file information. The trajectory file contains detailed information of all atoms at each time during the simulation, including but not limited to atom type, total charge, atomic coordinates, force, and system energy. By analyzing the trajectory file information, the high-precision acidity coefficient can be determined.

[0035] In some embodiments, the training data about the protein solution system corresponding to the protein can be obtained by quantum mechanics calculation. The training data includes total energy information of the protein solution system in a corresponding conformation, force information of each atom, etc. Based on the training data, the prediction information corresponding to the training data is obtained by the machine learning force field, and the prediction information includes total energy prediction information of the protein solution system in a corresponding conformation, force prediction information of each atom, etc. Based on the prediction information, the corresponding loss is calculated, and the parameters of the machine learning force field are updated in combination with an optimization algorithm (for example, Adam algorithm or its variants, Newton method, or quasi-Newton method, etc.), thereby obtaining a trained machine learning force field.

[0036] In some embodiments, the method further includes: step S14 (not shown), the device 1 constructs a quantum mechanics simulation model corresponding to the protein; based on the quantum mechanics simulation model, a corresponding non-equilibrium state data set is obtained by transition state scanning, wherein the non-equilibrium state data set includes structure information and corresponding system state information of the protein solution system corresponding to the protein in each reaction coordinate; based on the non-equilibrium state data set and the equilibrium state data set, the machine learning force field is trained and obtained, wherein the equilibrium state data set includes structure information and corresponding system state information of the equilibrium conformation corresponding to the protein solution system corresponding to the protein.

[0037] In some embodiments, the quantum mechanics simulation model comprises a quantum mechanics cluster (QM Cluster) model, or a quantum mechanics and molecular mechanics combined (QM / MM, Quantum Mechanics / Molecular Mechanics) model. For the quantum mechanics cluster model, the protein solution system corresponding to the protein is input into a molecular visualization software (e.g., PyMol, Chimera, or ChimeraX, etc.), and an active center is extracted using the molecular visualization software. The active center is usually selected as the atoms within a preset distance around the substrate in the system. For the quantum mechanics and molecular mechanics combined model, the QM and MM regions are distinguished using the molecular visualization software. The QM region is usually selected as the atoms within a preset distance around the substrate. The MM region comprises the remaining protein structure, solvent, ions, etc. after the QM region is divided. The aforementioned active center or structure with distinguished QM and MM regions is imported into GaussView (https: / / gaussian.com / ) for processing to obtain the corresponding quantum mechanics cluster model or quantum mechanics and molecular mechanics combined model. For example, the aforementioned active center or QM region structure is subjected to hydrogenation or chemical bond correction processing using GaussView, and the structure obtained by processing is used as the quantum mechanics simulation model for subsequent transition state scanning QM calculation. In some embodiments, the Nudged Elastic Band (NEB) algorithm or the Relaxed Scan algorithm can be used for transition state scanning. The process of the transition state scanning comprises: gradually changing the conformation corresponding to the protein solution system based on the selected reaction coordinate and the corresponding scanning information (e.g., scanning range, and scanning step, etc.), determining the atomic coordinates, force information, and total energy of the conformation corresponding to the reaction coordinate at each step through quantum mechanics calculation, and outputting the scanning data. The reaction coordinate is a variable used to describe the reaction process between biomolecules in the biological system, including but not limited to bond length, bond angle, or dihedral angle. Gaussian (https: / / gaussian.com / ), ORCA (https: / / www.faccts.de / orca / ), or PySCF (https: / / pyscf.org / ) quantum chemistry software can be used for quantum mechanics calculation in the aforementioned transition state scanning. The scanning data comprises atomic coordinate information, atomic force information, and total energy information of the conformation of the system at each reaction coordinate in the entire reaction process. That is, the scanning data comprises a large amount of non-equilibrium state atomic coordinate information, atomic force information, and total energy information data, which can be used as a non-equilibrium state data set. The system state information comprises total energy information of the system, and force information of the corresponding atoms.

[0038] In some embodiments, the equilibrium dataset can be obtained from an open source dataset or generated by itself. In some embodiments, the method further comprises: step S15 (not shown), based on the protein, the device 1 performs molecular dynamics simulation to obtain a plurality of frame information; using a molecular cutting algorithm, a plurality of structure information corresponding to each frame information is determined, and a quantum chemistry method is used to calculate the system state information corresponding to each structure information; based on the structure information and the corresponding system state information, a corresponding equilibrium dataset is determined. For example, through long-time molecular dynamics simulation, the system can traverse its phase space and reach equilibrium state, sufficient sampling is performed, and the plurality of frame information is obtained. Each frame information represents a snapshot of the system at a specific time, which includes but is not limited to the atomic number, the total charge number and the atomic coordinates of each atom in the system at the time corresponding to the frame information. In order to improve the sampling efficiency, enhanced sampling methods such as metadynamics or umbrella sampling can be used to deal with high energy barriers and long time scale problems in simulation, so as to obtain the frame information of the equilibrium state. For each frame information, a plurality of atoms in the system corresponding to the frame information are randomly selected as focus atoms. For each focus atom, based on the focus atom, a molecular cutting algorithm is used to determine a candidate structure information corresponding to the frame information. For example, a sphere is cut with the focus atom as the center and a preset distance threshold as the radius. Based on all atoms in the sphere, a corresponding atom set is established. It is determined whether the boundary atom in the atom set satisfies the preset cutting condition. The boundary atom exists a chemical bond to be cut, that is, the atom connected to the boundary atom through the cut chemical bond does not belong to the atom set. The preset cutting condition includes: the chemical bond to be cut of the boundary atom is a single bond; and the boundary atom is not a hydrogen atom. If all boundary atoms satisfy the preset cutting condition, the hydrogenation treatment is performed on the boundary atom based on the chemical bond to be cut of the boundary atom; the candidate structure information corresponding to the hydrogenation treated system is determined. If there is a boundary atom that does not satisfy the preset cutting condition, the atom connected to the boundary atom that does not satisfy the condition through the cut chemical bond is added to the atom set, so as to update the atom set. And continue to judge the preset cutting condition. For each candidate structure information, an analysis tool can be used to remove the candidate structure information with abnormalities (for example, the existence of a chemical bond with a bond length exceeding the corresponding bond length range, abnormal atom connection number, or open ring structure, etc.), to determine a plurality of structure information. Using a quantum chemistry method (perturbation theory, density functional theory, etc.), the system state information corresponding to each structure information is calculated. Further, the structure information and the corresponding system state information are used as an equilibrium dataset.

[0039] In some embodiments, the structural information is obtained by using a machine learning force field, and a system state prediction information is obtained. A mean squared error (MSE) loss function is used to calculate a corresponding loss, which includes but is not limited to energy loss and force loss. In combination with a corresponding optimization algorithm, the machine learning force field is updated by minimizing the loss. Here, the machine learning force field trained in combination with equilibrium and non-equilibrium data can more accurately capture proton transfer in simulation, so that the acidity coefficient can be more accurately determined by using the machine learning force field.

[0040] In some embodiments, the step S12 includes: the device 1 balances the simulation box; and a molecular dynamics simulation is performed on the balanced simulation box by using the trained machine learning force field to obtain corresponding trajectory file information. In some embodiments, in order to prevent the system from collapsing due to high energy at the beginning of the simulation, a steepest descent method or a conjugate gradient method can also be used to adjust the coordinates of each atom in the simulation box to minimize the energy for balancing. Then, a molecular dynamics simulation is performed on the balanced simulation box. Here, the balanced simulation box can be simulated for a long time (for example, at least 20 ns), so as to improve the reliability of the calculated acidity coefficient result.

[0041] In step S13, based on the trajectory file information, the acidity coefficient corresponding to each titratable site in the protein is determined. The titratable site is an atom in the protein that can be protonated / deprotonated. By analyzing the trajectory file information, the acidity coefficient corresponding to each titratable site in the protein can be calculated by using the Henderson-Hasselbalch equation.

[0042] In some embodiments, the step S13 includes: based on the trajectory file information, determining the protonation state concentration and the deprotonation state concentration corresponding to each titratable site in the protein; and based on the protonation state concentration and the deprotonation state concentration corresponding to the titratable site, and the pH information corresponding to the simulation box, determining the acidity coefficient corresponding to the titratable site. In some embodiments, based on the trajectory file information, the protonation concentration corresponding to the titratable site at each time of the molecular dynamics simulation can be determined, and then the average value thereof is determined as the protonation state concentration corresponding to the titratable site. The protonation concentration corresponding to the titratable site at each time can be determined based on the number of protonation states corresponding to the titratable site at the time. The calculation method of the protonation state concentration corresponding to the titratable site is as follows:

[0043]

[0044] wherein, is the concentration of the deprotonated state of the titratable site, is the number of the deprotonated state of the titratable site at a certain time of the molecular dynamics simulation (its value is 0 or 1), is the volume of the simulation box, is the Avogadro constant, is the average value of the protonated concentration at each time calculated in the bracket. The calculation of the concentration of the deprotonated state of the titratable site is the same as the calculation of the aforementioned protonated concentration, which is also based on the trajectory file information to determine the deprotonated concentration of the titratable site at each time of the molecular dynamics simulation, and then determine its average value as the deprotonated state concentration of the titratable site. The deprotonated concentration of the titratable site at each time can be determined based on the number of deprotonated states of the titratable site at that time. The calculation method of the deprotonated state concentration of the titratable site is:

[0045]

[0046] wherein, is the deprotonated state concentration of the titratable site, is the number of the deprotonated state of the titratable site at a certain time of the molecular dynamics simulation (its value is 0 or 1). Further, the Henderson-Hasselbalch equation can be used to calculate the acidity coefficient of the titratable site in combination with the pH information of the simulation box:

[0047]

[0048] Since the simulation box is added with a buffer to maintain the stability of the pH environment in the simulation box, the calculation of the acidity coefficient can be directly based on the pH set by the simulation box.

[0049] In some embodiments, in order to make the calculated acidity coefficient more accurate, the actual pH information of the simulation box can also be calculated based on the simulation situation. The step S13 further comprises: determining the hydrogen ion concentration corresponding to the simulation box based on the trajectory file information; determining the pH information corresponding to the simulation box based on the hydrogen ion concentration. In some embodiments, the number of hydrogen ions at each time of the molecular dynamics simulation can be determined based on the trajectory file information, and then the hydrogen ion concentration at each time is determined. By averaging the hydrogen ion concentration at each time, the hydrogen ion concentration corresponding to the simulation box is determined, and then the pH information corresponding to the simulation box is determined. The calculation method is:

[0050]

[0051] wherein, C(t) is the hydrogen ion concentration corresponding to the simulation box at time t, N(t) is the number of hydrogen ions at time t in the molecular dynamics simulation, C(t) is the hydrogen ion concentration corresponding to the simulation box at time t,

[0052] Figure 2 Fig. 1 shows a device structure for determining acidity coefficients according to an embodiment of the present application, which comprises a first module 11, a second module 12 and a third module 13. The first module 11 constructs a simulation box corresponding to a protein, wherein the simulation box comprises the protein, a corresponding solvent and a buffer; the second module 12 performs molecular dynamics simulation on the simulation box using a trained machine learning force field to obtain corresponding trajectory file information; and the third module 13 determines acidity coefficients corresponding to each titratable site in the protein based on the trajectory file information. Here, the first module 11, the second module 12 and the third module 13 correspond to the specific embodiments of the aforementioned steps S11, S12 and S13 respectively, and thus will not be described again, but are included herein by reference. Figure 2 Fig. 1 shows a device structure for determining acidity coefficients according to an embodiment of the present application, which comprises a first module 11, a second module 12 and a third module 13. The first module 11 constructs a simulation box corresponding to a protein, wherein the simulation box comprises the protein, a corresponding solvent and a buffer; the second module 12 performs molecular dynamics simulation on the simulation box using a trained machine learning force field to obtain corresponding trajectory file information; and the third module 13 determines acidity coefficients corresponding to each titratable site in the protein based on the trajectory file information. Here, the first module 11, the second module 12 and the third module 13 correspond to the specific embodiments of the aforementioned steps S11, S12 and S13 respectively, and thus will not be described again, but are included herein by reference.

[0053] In some embodiments, the first module 11 comprises a first unit 111 (not shown) and a second unit 112 (not shown). The first unit 111 places the protein in the simulation box based on structure information corresponding to the protein and adds a corresponding solvent; and the second unit 112 replaces part of the solvent molecules added in the simulation box with buffer molecules based on the volume of the simulation box and a preset buffer concentration. Here, the specific embodiments of the first unit 111 and the second unit 112 are the same as or similar to the specific embodiments of the aforementioned steps S111 and S112 respectively, and thus will not be described again, but are included herein by reference.

[0054] In some embodiments, the device 1 further comprises a fourth module 14 (not shown). The fourth module 14 constructs a quantum mechanics simulation model corresponding to the protein; obtains a corresponding non-equilibrium state dataset through transition state scanning based on the quantum mechanics simulation model, wherein the non-equilibrium state dataset comprises structure information and corresponding system state information of a protein solution system corresponding to the protein at each reaction coordinate; and trains the machine learning force field based on the non-equilibrium state dataset and an equilibrium state dataset, wherein the equilibrium state dataset comprises structure information and corresponding system state information of an equilibrium conformation corresponding to the protein solution system of the protein. Here, the specific embodiment of the fourth module 14 is the same as or similar to the specific embodiment of the aforementioned step S14, and thus will not be described again, but is included herein by reference.

[0055] In some embodiments, the device 1 further comprises a fifth module 15 (not shown). The fifth module 15 performs molecular dynamics simulation based on the protein to obtain a plurality of frame information; determines a plurality of structure information corresponding to each frame information by using a molecular cutting algorithm, and calculates a system state information corresponding to each structure information by using a quantum chemistry method; and determines a corresponding equilibrium data set based on the structure information and the corresponding system state information. Here, the specific implementation of the fifth module 15 is the same as or similar to that of the foregoing embodiment of the step S15, and thus will not be described herein again by way of reference.

[0056] Figure 3 An example system that can be used to implement various embodiments described herein is shown. As Figure 3 As shown, in some embodiments, the system 300 can function as any of the devices described herein. In some embodiments, the system 300 can include one or more computer-readable media (e.g., system memory or NVM / storage 320) having instructions and one or more processors (e.g., processor(s) 305) coupled to the one or more computer-readable media and configured to execute the instructions to implement modules to perform the actions described herein.

[0057] For one embodiment, the system control module 310 can include any suitable interface controllers to provide for any suitable interface to at least one of the processor(s) 305 and / or any suitable device or component in communication with the system control module 310.

[0058] The system control module 310 can include a memory controller module 330 to provide an interface to the system memory 315. The memory controller module 330 can be a hardware module, a software module, and / or a firmware module.

[0059] The system memory 315 can be used, for example, to load and store data and / or instructions for the system 300. For one embodiment, the system memory 315 can include any suitable volatile memory, such as suitable DRAM. In some embodiments, the system memory 315 can include double data rate type four synchronous dynamic random access memory (DDR4 SDRAM).

[0060] For one embodiment, the system control module 310 can include one or more input / output (I / O) controller(s) to provide an interface to the NVM / storage 320 and the communication interface(s) 325.

[0061] For example, NVM / storage 320 can be used to store data and / or instructions. NVM / storage 320 can include any suitable non-volatile memory (e.g., flash memory) and / or can include any suitable non-volatile storage device(s) (e.g., hard disk drive(s) (HDD(s)), compact disk (CD) drive(s), and / or digital versatile disk (DVD) drive(s)).

[0062] NVM / storage 320 can include a storage resource that is physically part of the device on which system 300 is installed or that is accessed via the device but that is not necessarily part of the device. For example, NVM / storage 320 can be accessed over a network via communication interface(s) 325.

[0063] Communication interface(s) 325 can provide an interface for system 300 to communicate with one or more networks and / or any other suitable device. System 300 can communicate wirelessly with one or more components of a wireless network according to any of one or more wireless network standards and / or protocols.

[0064] For one embodiment, at least one of processor(s) 305 can be packaged together with logic of one or more controllers of system control module 310 (e.g., memory controller module 330). For one embodiment, at least one of processor(s) 305 can be packaged together with logic of one or more controllers of system control module 310 to form a system-in-a-package (SiP). For one embodiment, at least one of processor(s) 305 can be integrated on the same die with logic of one or more controllers of system control module 310. For one embodiment, at least one of processor(s) 305 can be integrated on the same die with logic of one or more controllers of system control module 310 to form a system-on-a-chip (SoC).

[0065] In various embodiments, system 300 can be, but is not limited to, a server, a workstation, a desktop computing device, or a mobile computing device (e.g., a laptop computing device, a handheld computing device, a tablet, a netbook, etc.). In various embodiments, system 300 can have more or less components, and / or different architectures. For example, in some embodiments, system 300 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including touch screen displays), non-volatile memory port, multiple antennas, a graphics chip, an application specific integrated circuit (ASIC), and a speaker.

[0066] In addition to the methods and devices described in the above embodiments, the present application also provides a computer readable storage medium storing computer code which, when executed, performs the method of any preceding item.

[0067] The present application also provides a computer program product which, when executed by a computer device, performs the method of any preceding item.

[0068] The present application also provides a computer device comprising:

[0069] one or more processors;

[0070] a memory for storing one or more computer programs;

[0071] when the one or more computer programs are executed by the one or more processors, the one or more processors implement the method of any preceding item.

[0072] It should be noted that the present application can be implemented in software and / or a combination of software and hardware, for example, using an application specific integrated circuit (ASIC), a general purpose computer or any other similar hardware device. In one embodiment, the software program of the present application can be executed by a processor to implement the steps or functions described above. Similarly, the software program of the present application (including related data structures) can be stored in a computer readable recording medium, such as a RAM memory, a magnetic or optical drive or a soft disk and the like. In addition, some steps or functions of the present application can be implemented in hardware, for example, as a circuit cooperating with the processor to perform the respective steps or functions.

[0073] In addition, part of the present application can be applied as a computer program product, for example, computer program instructions, when executed by a computer, through the operation of the computer, the method and / or technical solutions according to the present application can be invoked or provided. Those skilled in the art should understand that the form of computer program instructions in computer readable medium includes but is not limited to source file, executable file, installation package file and the like, and accordingly, the way of computer program instructions executed by computer includes but is not limited to: the computer directly executes the instructions, or the computer compiles the instructions and then executes the corresponding compiled program, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Here, the computer readable medium can be any available computer readable storage medium or communication medium accessible to the computer.

[0074] Communication media includes wired and wireless media involving the transmission of signals, including computer readable instructions, data structures, program modules, or other data, between or among computers, mobile devices, and communications systems over the aforementioned media. Modulated data signals can be passed through wired and / or wireless transmission mediums, including those inherently connected to a system, such as over a bus, or those externally connected, such as through electrical cable or optical fiber. Modulated data signals can be passed over a carrier wave, in baseband, or use some combination of a carrier wave and baseband. The known uses of "modulated data signal" and / or "carrier wave" include, but are not limited to, those uses disclosed in U.S. Patent No. 5,452,345, which is incorporated by reference as if fully set forth herein.

[0075] By way of example, and not limitation, computer readable storage media can include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. For example, computer readable storage media includes, but is not limited to, RAM, such as SRAM, DRAM, or RAM; ROM, such as PROM, EPROM, EEPROM, or Flash memory; magnetic and ferromagnetic / ferroelectric storage, such as MRAM, FeRAM; and magnetic and optical storage, such as hard disks, tapes, CDs, DVDs, or other media now known or later developed that are used to store computer readable information / data for use in computers.

[0076] Herein, according to one embodiment of the present application includes an apparatus comprising a memory for storing computer program instructions and a processor for executing the program instructions, wherein when the computer program instructions are executed by the processor, the apparatus is triggered to run the method and / or technical solutions based on the aforementioned embodiments according to the present application.

[0077] It will be obvious to a person skilled in the art that the application is not limited to the details of the above-described exemplary embodiments but can be implemented in other embodiments without departing from the scope of the application. The application is therefore not limited by the described examples but can vary within the scope of the claims and their equivalents. No reference signs in the claims should be considered as limiting the scope of the claims. Furthermore, it is obvious that the wording "comprising" does not exclude other parts and does not exclude other steps. Singularity is not excluded with respect to plurality and vice versa. Multiple units or devices recited in a device claim can also be implemented by one unit or device by means of software or hardware. The terms first, second and the like do not denote any ordering, but rather serve as names.

Claims

1. A method for determining an acidity factor, wherein, The method comprises: constructing a simulation box corresponding to the protein, wherein the simulation box comprises the protein, a corresponding solvent and a buffer; performing molecular dynamics simulation on the simulation box by using a trained machine learning force field to obtain corresponding trajectory file information; determining the acidity coefficients of each titratable site in the protein based on the trajectory file information.

2. The method of claim 1, wherein, The constructing a simulation box corresponding to the protein, wherein the simulation box comprises the protein, a corresponding solvent and a buffer comprises: placing the protein in the simulation box based on the structure information corresponding to the protein, and adding a corresponding solvent; replacing part of the solvent molecules in the added solvent in the simulation box with buffer molecules based on the volume of the simulation box and a preset buffer concentration.

3. The method of claim 2, wherein, The replacing part of the solvent molecules in the added solvent in the simulation box with buffer molecules based on the volume of the simulation box and a preset buffer concentration comprises: i randomly selecting one solvent molecule from the added solvent in the simulation box and replacing it with a buffer molecule corresponding to the buffer; determining whether the buffer molecule satisfies a replacement condition, if the condition is not satisfied, canceling the replacement, restoring the buffer molecule to a solvent molecule, and returning to step i; otherwise, continuing to perform step i until the number of added buffer molecules in the simulation box reaches a set number, wherein the set number is determined based on the volume of the simulation box and a preset buffer concentration.

4. The method of claim 3, wherein, The replacement condition comprises: the minimum distance between the buffer molecule and the protein is greater than or equal to a corresponding distance threshold; and the buffer molecule is located inside the simulation box.

5. The method of any one of claims 1 to 4, wherein, The performing molecular dynamics simulation on the simulation box by using a trained machine learning force field to obtain corresponding trajectory file information comprises: equilibrating the simulation box; performing molecular dynamics simulation on the equilibrated simulation box by using a trained machine learning force field to obtain corresponding trajectory file information.

6. The method of any one of claims 1 to 5, wherein, The method further comprises: constructing a quantum mechanics simulation model corresponding to the protein; obtaining a non-equilibrium state data set by transition state scanning based on the quantum mechanics simulation model, wherein the non-equilibrium state data set comprises structure information and corresponding system state information of a protein solution system corresponding to the protein at each reaction coordinate; training the machine learning force field based on the non-equilibrium state data set and an equilibrium state data set, wherein the equilibrium state data set comprises structure information and corresponding system state information of an equilibrium conformation of the protein solution system corresponding to the protein.

7. The method of claim 6, wherein, The method further comprises: performing molecular dynamics simulation based on the protein to obtain corresponding multiple frame information; determining multiple structure information corresponding to each frame information by using a molecular cutting algorithm, and calculating system state information corresponding to each structure information by using quantum chemistry method; determining a corresponding equilibrium state data set based on the structure information and the corresponding system state information.

8. The method of any one of claims 1 to 7, wherein, The determining the acidity coefficients of each titratable site in the protein based on the trajectory file information comprises: determine, based on the trajectory file information, protonated state concentration and deprotonated state concentration corresponding to each titratable site in the protein; determine, based on the protonated state concentration and deprotonated state concentration corresponding to each titratable site, in combination with the p H information corresponding to the simulation box, an acidity coefficient corresponding to each titratable site.

9. The method of claim 8, wherein, The determining, based on the trajectory file information, of the acidity coefficient corresponding to each titratable site in the protein further comprises: determine, based on the trajectory file information, a hydrogen ion concentration corresponding to the simulation box; determine, based on the hydrogen ion concentration, p H information corresponding to the simulation box.

10. A computer device for determining an acidity coefficient, comprising a memory, a processor and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 9.

11. A computer readable storage medium having stored thereon computer programs / instructions, characterized in that, The computer program / instructions, when executed by the processor, implement the steps of the method according to any one of claims 1 to 9.

12. A computer program product comprising a computer program, characterized in that, The computer program, when executed by the processor, implements the steps of the method according to any one of claims 1 to 9.

Citation Information

Cited By

  • Method, equipment, medium and program product for constructing machine learning force field for non-equilibrium state biological system simulation

    CN121096416A