Structure extraction program, structure extraction method, and information processing device
The structure extraction program enhances molecular simulation analysis by prioritizing samples reflecting frequent intermolecular interactions through a score function based on distance thresholds, improving the understanding of simulation outcomes.
Patent Information
- Application Number
- JP2024517700
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-04-27
- Publication Date
- 2025-10-16
- Estimated Expiration
- 2042-04-27
AI Technical Summary
Molecular simulations generate numerous molecular configuration samples, including those that do not reflect frequently occurring intermolecular interactions, making it difficult to understand the simulation results effectively.
A structure extraction program that calculates a score function based on a distance threshold to determine the priority of molecular configuration samples, allowing extraction of those that reflect frequently occurring intermolecular interactions.
Enables the extraction of molecular configuration samples that are more useful for understanding intermolecular interactions, improving the accuracy and relevance of simulation results.
Smart Images

Figure 0007755205000007 
Figure 0007755205000008 
Figure 0007755205000009
Abstract
Description
[Technical Field]
[0001] The present invention relates to a structure extraction program, a structure extraction method, and an information processing device. [Background technology]
[0002] Molecular simulation is one type of computer simulation. Molecular simulation is sometimes used to analyze intermolecular interactions, which show the forces of attraction between different molecules. Analysis of intermolecular interactions can be applied in fields such as life science and materials development.
[0003] For example, molecular dynamics simulation calculates the interaction force on each atom based on the equation of motion while advancing time in small time increments, and iteratively updates the position and velocity of each atom. Also, for example, Monte Carlo simulation randomly generates molecular configuration samples according to a probability distribution based on the Boltzmann factor, and calculates expected physical properties from multiple molecular configuration samples. The Boltzmann factor indicates a weight that is inversely proportional to the energy of the molecular configuration.
[0004] A simulation device has been proposed that uses molecular dynamics simulation to reproduce chemical changes involving the rearrangement of covalent bonds in large-scale atomic ensembles. A structure determination method has also been proposed that generates three-dimensional structures of dynamic organic molecules with interatomic bonds with variable angles and simulates the changes in the three-dimensional structure in a population of dynamic organic molecules. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Japanese Patent Application Laid-Open No. 2006-190234 [Patent Document 2] International Publication No. 2009 / 034297 Summary of the Invention [Problem to be solved by the invention]
[0006] A molecular simulation generates a plurality of molecular configuration samples during the process. For example, a molecular dynamics simulation generates a plurality of molecular configuration samples corresponding to a plurality of time points. Furthermore, for example, a Monte Carlo simulation randomly generates a plurality of molecular configuration samples. To assist in understanding the simulation results, an information processing device may wish to visualize and display some of the molecular configuration samples (for example, any one of the molecular configuration samples).
[0007] In this case, it is important to select which molecular configuration samples to extract from the multiple molecular configuration samples generated. The molecular configuration samples generated during the simulation process may include molecular configuration samples that do not reflect frequently occurring intermolecular interactions. For example, even if a group of atoms in one molecule and a group of atoms in another molecule are close to each other in a molecular configuration sample, this may be due to chance proximity and may not be based on frequently occurring intermolecular interactions. Such molecular configuration samples may be less useful for understanding the simulation results.
[0008] Therefore, in one aspect, the present invention aims to extract molecular configuration samples that reflect frequently occurring intermolecular interactions from the results of molecular simulations. [Means for solving the problem]
[0009] In one embodiment, a structure extraction program is provided that causes a computer to execute the following processes: acquire a plurality of molecular configuration samples, each containing positional information for a plurality of atoms; generate a score function, which is a function for calculating weights for pairs of atoms among the plurality of atoms that belong to different molecules and includes a parameter indicating a distance threshold, based on the distances between the pairs of atoms in the plurality of molecular configuration samples; determine a specific value to be set for the parameter based on a change in the weight calculated by the score function in response to a change in the value of the parameter; and determine the priority of the plurality of molecular configuration samples based on the specific weight calculated by the score function using the specific value and the distance between the pairs of atoms in each of the plurality of molecular configuration samples.
[0010] In one aspect, a structure extraction method executed by a computer is provided.In another aspect, an information processing device having a storage unit and a processing unit is provided. [Effects of the Invention]
[0011] In one aspect, molecular configuration samples that reflect frequently occurring intermolecular interactions can be extracted from the results of molecular simulations. The above and other objects, features and advantages of the present invention will become apparent from the following description taken in conjunction with the accompanying drawings illustrating preferred embodiments of the present invention. [Brief explanation of the drawings]
[0012] [Figure 1] FIG. 1 is a diagram illustrating an information processing apparatus according to a first embodiment. [Figure 2] FIG. 2 is a block diagram illustrating an example of hardware of an information processing device. [Figure 3] FIG. 10 is a diagram showing examples of a plurality of molecular arrangement samples. [Figure 4] FIG. 10 is a diagram illustrating an example of statistical information of physical property values. [Figure 5] FIG. 10 is a diagram showing examples of interatomic distances and cutoff distances. [Figure 6] 10 is a graph showing an example of the derivative of the average atomic density. [Figure 7] FIG. 10 is a diagram showing a comparison between experimental results and simulation results. [Figure 8] FIG. 2 is a block diagram illustrating an example of functions of the information processing device. [Figure 9] 10A and 10B are diagrams showing examples of a molecule list and a molecule configuration table. [Figure 10] FIG. 10 is a diagram illustrating an example of a score function table. [Figure 11] 10 is a flowchart showing an example of a procedure for visualizing a molecular arrangement. DETAILED DESCRIPTION OF THE INVENTION
[0013] The present embodiment will be described below with reference to the drawings. [First embodiment] A first embodiment will be described.
[0014] FIG. 1 is a diagram illustrating an information processing apparatus according to a first embodiment. The information processing device 10 of the first embodiment extracts some molecular configuration samples from a plurality of molecular configuration samples generated by molecular simulation. The molecular simulation may be executed by the information processing device 10 or another information processing device. The information processing device 10 may be a client device or a server device. The information processing device 10 may also be called a computer, a structure extraction device, or a simulation device.
[0015] The information processing device 10 has a storage unit 11 and a processing unit 12. The storage unit 11 may be a volatile semiconductor memory such as a random access memory (RAM), or a non-volatile storage such as a hard disk drive (HDD) or a flash memory. The processing unit 12 is a processor such as a central processing unit (CPU), a graphics processing unit (GPU), or a digital signal processor (DSP). However, the processing unit 12 may also include an electronic circuit such as an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA). The processor executes a program stored in a memory such as a RAM (which may be the storage unit 11). A set of processors may be called a multiprocessor or simply a "processor."
[0016] The memory unit 11 stores a plurality of molecular configuration samples, including molecular configuration samples 13 and 14. The molecular configuration samples 13 and 14 indicate the relative positional relationship between two or more molecules, i.e., the mutual arrangement of two or more molecules. The molecular configuration samples 13 and 14 may indicate the molecular configuration in two-dimensional space or in three-dimensional space. The molecular configuration samples 13 and 14 may be referred to as "structures" or "three-dimensional structures." The molecular configuration samples 13 and 14 include position information for each of a plurality of atoms included in two or more molecules. The position information may include two-dimensional coordinates or three-dimensional coordinates.
[0017] The molecular configuration samples 13 and 14 are generated through a molecular simulation that analyzes intermolecular interactions. The intermolecular interactions represent the intermolecular forces that attract different molecules to each other. The molecular simulation may be executed by the information processing device 10 or by another information processing device. The molecular simulation may be, for example, a molecular dynamics simulation or a Monte Carlo simulation. In the case of a molecular dynamics simulation, the molecular configuration samples 13 and 14 represent molecular configurations at different times. In the case of a Monte Carlo simulation, the molecular configuration samples 13 and 14 represent different molecular configurations that are randomly generated.
[0018] The processing unit 12 generates a score function 15 using the molecular configuration samples 13 and 14. The score function 15 calculates weights for pairs of atoms belonging to different molecules. The atoms forming the pairs may be some of the atoms included in the molecular configuration samples 13 and 14. For example, some elements such as hydrogen may be excluded from the pairs. The score function 15 includes a parameter 16 that indicates a distance threshold. The parameter 16 may also be called a cutoff distance or cutoff radius. At this point, the value of the parameter 16 has not yet been optimized. If the value of the parameter 16 changes, the weight calculated by the score function 15 also changes.
[0019] The processing unit 12 generates a score function 15 based on the distance between pairs of atoms in the molecular configuration samples 13 and 14. The distance between a certain pair varies depending on the molecular configuration sample. The score function 15 calculates a large weight for a pair of atoms that are stably close to each other. For example, the score function 15 calculates a weight for a certain pair according to the proportion of molecular configuration samples in which the distance between the pair is equal to or less than the value of the parameter 16. The proportion may be the weight.
[0020] Furthermore, for example, the processing unit 12 calculates the value of a potential function from each molecular configuration sample for a certain pair. The potential function indicates the magnitude of the interaction between atoms. The potential function may depend on the distance between the atoms, and the value of the potential function may decrease as the distance decreases. The potential function may also depend on the elements of the two atoms, and the value of the potential function may increase when the two atoms are carbon atoms. The score function 15 calculates a weight according to the average of the potential function values. The average may be the weight. However, for molecular configuration samples where the distance between the pair exceeds the value of the parameter 16, the value of the potential function is considered to be zero.
[0021] Parameter 16 may be used to determine whether two atoms are neighbors or not, and may also be used to limit the range of the neighboring region to be considered in calculating the weight. For example, if the distance between two atoms exceeds the value of parameter 16, the two atoms are determined not to be neighbors. Also, for example, mutual configurations in the molecular configuration sample whose distance exceeds the value of parameter 16 are not considered in calculating the weights of the two atoms.
[0022] For example, the first atom, the second atom, and the third atom are not hydrogen atoms and belong to different molecules. In the molecular configuration sample 13, the distance between the first atom and the second atom is R 1 12 and the distance between the first and third atoms is R 1 13 R 1 12 is greater than R0, and R 1 13 In molecular configuration sample 14, the distance between the first atom and the second atom is R 2 12 and the distance between the first and third atoms is R 2 13 R 2 12 is less than or equal to R0, and R 2 13 is less than or equal to R0.
[0023] When R0 is set to the parameter 16, the score function 15 calculates the R 2 12 The score function 15 calculates a weight taking into account the R 1 13 and R of molecular configuration sample 14 2 13 Calculate the weight taking into account the above.
[0024] The processing unit 12 determines a specific value to be set for the parameter 16 based on a change in the weight of the score function 15 in response to a change in the value of the parameter 16. For example, the processing unit 12 calculates an evaluation value by dividing the weight of the score function 15 by a space size corresponding to the value of the parameter 16, and monitors the change in the evaluation value in response to a change in the value of the parameter 16. When the molecular configuration samples 13 and 14 represent molecular configurations in a three-dimensional space, the space size may be the cube of the value of the parameter 16. When there are multiple pairs of atoms, the evaluation value may be the sum of the weights of the multiple pairs divided by the space size. The processing unit 12 may determine a specific value that maximizes the differentiation of the evaluation value with respect to the value of the parameter 16, i.e., maximizes the unit change in the evaluation value.
[0025] The processing unit 12 sets the determined specific value, i.e., the determined distance threshold, to the parameter 16. The processing unit 12 determines priorities 17 of the multiple molecular configuration samples including the molecular configuration samples 13 and 14 based on specific weights calculated by the score function 15 using the specific value and the distances between pairs of atoms in each molecular configuration sample.
[0026] For example, the processing unit 12 calculates a score for each molecular configuration sample based on whether the distance between the atom pairs is equal to or less than the value of the parameter 16 and the weight calculated by the score function 15, and determines the priority order 17 based on the score. The processing unit 12 may calculate the score using the weight of the atom pairs whose distance is equal to or less than the value of the parameter 16. When there are multiple atom pairs, the score may be the sum of the weights of the pairs whose distance is equal to or less than the value of the parameter 16. The processing unit 12 may rank the multiple molecular configuration samples in descending order of score. Note that only some of the molecular configuration samples may be ranked.
[0027] The processing unit 12 may generate image data that visualizes the position information of atoms included in some molecular configuration samples selected based on the priority order 17 (for example, one or more molecular configuration samples selected in descending order of priority). In the image data, the position of each atom may be represented by a point, a circle, or a sphere. The processing unit 12 may output the generated image data. For example, the image data may be displayed on a display device, may be stored in nonvolatile storage, or may be transmitted to another information processing device. Furthermore, the visualization of the molecular configuration may be performed by another information processing device.
[0028] As described above, the information processing device 10 of the first embodiment generates a score function 15 including a parameter 16 indicating a distance threshold based on the distance between atom pairs in a plurality of molecular configuration samples. The information processing device 10 determines a specific value to be set for the parameter 16 based on a change in the weight calculated by the score function 15 in response to a change in the value of the parameter 16. The information processing device 10 determines a priority order 17 for a plurality of molecular configuration samples based on the weight calculated by the score function 15 and the distance between atom pairs in each molecular configuration sample.
[0029] This allows molecular configuration samples that reflect frequently occurring intermolecular interactions to be extracted from multiple molecular configuration samples. Therefore, compared to other extraction methods, such as extracting the last generated molecular configuration sample, molecular configuration samples that are more useful for understanding the analysis results of intermolecular interactions obtained by molecular simulation are extracted. Furthermore, based on the multiple molecular configuration samples, a score function 15 is generated and a distance threshold is determined. Therefore, an appropriate score function 15 that is suited to the current molecular simulation is generated.
[0030] For example, if the distance threshold is too small, the score function 15 may not sufficiently cover atoms that are stably present nearby. On the other hand, if the distance threshold is too large, the score function 15 may take into account interactions with distant atoms that do not contribute to a stable molecular configuration, which may reduce the accuracy of the weights of the score function 15. In this regard, the information processing device 10 can automatically adjust the distance threshold.
[0031] Furthermore, for example, the magnitude of the interaction and the direction of attraction between two molecules may differ depending on the molecules being simulated. Also, simulation conditions such as surrounding water molecules may differ for each molecular simulation. In this regard, the information processing device 10 can calculate a distance threshold appropriate for the current molecular simulation.
[0032] Furthermore, the information processing device 10 may generate and output image data that visualizes the positional information of atoms included in the molecular configuration sample with the highest priority 17. At this time, the information processing device 10 may output the image data together with statistical information on physical property values generated from a plurality of molecular configuration samples. This provides a visualization result that is consistent with the statistical information. Furthermore, the information processing device 10 may generate the score function 15 so as to calculate a weight according to the proportion of molecular configuration samples in which the distance between an atom pair is equal to or less than the value of the parameter 16. This allows the score function 15 to evaluate whether a certain atom pair is stably located nearby.
[0033] Furthermore, the information processing device 10 may monitor changes in an evaluation value obtained by dividing the weight of the score function 15 by a space size corresponding to the value of the parameter 16 in response to changes in the value of the parameter 16, and determine the value of the parameter 16 so that the derivative of the evaluation value is maximized. This allows the information processing device 10 to determine a threshold value that covers a range in which there are many atoms having stable interactions. Furthermore, the information processing device 10 may calculate a score for each molecular configuration sample based on whether the distance between the atom pair is equal to or less than a threshold value and the weight of the score function 15, and determine a priority order 17 based on the score. This makes it easy to identify molecular configuration samples that reflect frequently occurring intermolecular interactions.
[0034] [Second embodiment] Next, a second embodiment will be described. The information processing device 100 of the second embodiment analyzes intermolecular interactions by executing molecular simulations and outputs the analysis results. For example, the information processing device 100 analyzes intermolecular interactions between proteins and low-molecular-weight compounds to support drug discovery. The molecular simulations include molecular dynamics simulations and Monte Carlo simulations. The analysis results include statistical information indicating physical property values such as interaction energy. The analysis results also include image data that visualizes typical molecular configuration samples. The typical molecular configuration samples reflect typical intermolecular interactions indicated by the statistical information.
[0035] The information processing device 100 may be a client device or a server device. The information processing device 100 may be called a computer, a structure extraction device, or a simulation device. The information processing device 100 may execute a structure extraction program. The information processing device 100 corresponds to the information processing device 10 of the first embodiment.
[0036] FIG. 2 is a block diagram illustrating an example of hardware of an information processing device. The information processing device 100 has a CPU 101, a RAM 102, a HDD 103, a GPU 104, an input interface 105, a medium reader 106, and a communication interface 107, all connected via a bus. The CPU 101 corresponds to the processing unit 12 in the first embodiment. The RAM 102 or the HDD 103 corresponds to the storage unit 11 in the first embodiment.
[0037] The CPU 101 is a processor that executes program instructions. The CPU 101 loads programs and data stored in the HDD 103 into the RAM 102 and executes the programs. The information processing device 100 may have multiple processors.
[0038] The RAM 102 is a volatile semiconductor memory that temporarily stores programs executed by the CPU 101 and data used in calculations by the CPU 101. The information processing device 100 may have a type of volatile memory other than a RAM.
[0039] The HDD 103 is a nonvolatile storage that stores software programs such as an operating system (OS), middleware, and application software, as well as data. The information processing device 100 may also have other types of nonvolatile storage, such as a flash memory or an SSD (Solid State Drive).
[0040] The GPU 104 performs image processing in cooperation with the CPU 101 and outputs an image to a display device 111 connected to the information processing device 100. The display device 111 is, for example, a CRT (Cathode Ray Tube) display, a liquid crystal display, an organic EL (Electro Luminescence) display, or a projector. Other types of output devices, such as a printer, may be connected to the information processing device 100. The GPU 104 may also be used as a GPGPU (General Purpose Computing on Graphics Processing Unit). The GPU 104 may execute a program in response to an instruction from the CPU 101. The information processing device 100 may have a volatile semiconductor memory other than the RAM 102 as a GPU memory.
[0041] The input interface 105 receives an input signal from an input device 112 connected to the information processing device 100. The input device 112 is, for example, a mouse, a touch panel, or a keyboard. A plurality of input devices may be connected to the information processing device 100.
[0042] The medium reader 106 is a reading device that reads programs and data recorded on the recording medium 113. The recording medium 113 is, for example, a magnetic disk, an optical disk, or a semiconductor memory. Magnetic disks include flexible disks (FDs) and HDDs. Optical disks include compact discs (CDs) and digital versatile discs (DVDs). The medium reader 106 copies the programs and data read from the recording medium 113 to other recording media such as the RAM 102 or the HDD 103. The read programs may be executed by the CPU 101.
[0043] The recording medium 113 may be a portable recording medium. The recording medium 113 may be used to distribute programs and data. The recording medium 113 and the HDD 103 may also be referred to as computer-readable recording media.
[0044] The communication interface 107 communicates with other information processing devices via the network 114. The communication interface 107 may be a wired communication interface connected to a wired communication device such as a switch or a router, or may be a wireless communication interface connected to a wireless communication device such as a base station or an access point.
[0045] Next, molecular simulation will be described. The information processing device 100 performs molecular simulation. Molecular simulation can be applied to various fields such as life science, material development, and drug discovery support. Molecular simulation for drug discovery support may analyze intermolecular interactions between target proteins and low-molecular-weight compounds, or antigen-antibody reactions. In the following explanation, it is assumed that the information processing device 100 analyzes intermolecular interactions between proteins and low-molecular-weight compounds.
[0046] Molecular simulations such as molecular dynamics simulations and Monte Carlo simulations generate a large number of molecular configuration samples in which two or more molecules are arranged differently relative to one another. For example, several thousand to several hundred thousand molecular configuration samples are generated. Since the simulation space in the second embodiment is three-dimensional, the molecular configuration samples may also be referred to as three-dimensional structures. The information processing device 100 generates statistical information on physical property values from these multiple molecular configuration samples. An example of the statistical information is the interaction energy between a protein and a low-molecular-weight compound.
[0047] The molecular dynamics simulation receives information indicating the elements of the multiple atoms that form the molecule, and information indicating the initial positions and initial velocities of the multiple atoms. The molecular dynamics simulation calculates the interaction force that each atom experiences based on the current molecular configuration, and calculates the position and velocity of each atom after a short time according to the equation of motion shown in formula (1). In formula (1), F i is the interaction force experienced by atom i, m i is the mass of atom i, r iis the position of atom i, and t is the time. The molecular dynamics simulation iteratively updates the position and velocity of each atom as time progresses. Therefore, the molecular dynamics simulation generates multiple molecular configuration samples corresponding to multiple times.
[0048]
number
[0049] Monte Carlo simulation repeats random generation of molecular configuration samples according to the probability distribution indicated by the Boltzmann factor in formula (2). In formula (2), E is the energy of the atomic configuration, k B is the Boltzmann constant, and T is the temperature. According to the probability distribution indicated by the Boltzmann factor, the probability of generating a stable atomic configuration with a small energy E is high, and the probability of generating an unstable atomic configuration with a large energy E is low. Therefore, Monte Carlo simulation generates multiple molecular configuration samples corresponding to multiple trials.
[0050]
number
[0051] FIG. 3 is a diagram showing examples of a plurality of molecular arrangement samples. Molecular configuration samples 31, 32, and 33 are molecular configuration samples corresponding to different times generated by molecular dynamics simulation. Molecular configuration samples 31, 32, and 33 each show the relative position of a low molecular weight compound 35 with respect to a protein 34. However, in addition to the protein 34 and the low molecular weight compound 35, other molecules such as water molecules are also present. The relative positions of the protein 34 and the low molecular weight compound 35 change over time. Therefore, the positions of the atoms that form the molecules differ between molecular configuration samples 31, 32, and 33.
[0052] FIG. 4 is a diagram showing an example of statistical information of physical property values. Graph 36 shows physical property values calculated from molecular configuration samples generated by molecular dynamics simulation. The horizontal axis of graph 36 represents time (units: nanoseconds). The vertical axis of graph 36 represents the interaction energy (units: kilojoules / mole) between protein 34 and low molecular weight compound 35. The interaction energy is calculated from the relative position of low molecular weight compound 35 with respect to protein 34 according to an equation that represents the laws of physics. Over time, the mutual configuration of protein 34 and low molecular weight compound 35 changes, causing the interaction energy to fluctuate. The information processing device 100 is able to calculate the average energy from graph 36. The statistical information may be the interaction energy at each time or the average energy.
[0053] The information processing device 100 outputs the above-described statistical information as the analysis result of the intermolecular interactions. For example, the information processing device 100 displays the statistical information on the display device 111. At this time, the information processing device 100 outputs image data that visualizes one or a small number of molecular arrangement samples together with the statistical information to help understand the statistical information. The image data is visual information that represents the positions of atoms indicated by the molecular arrangement sample using figures such as circles. For example, the information processing device 100 displays this image data on the display device 111.
[0054] It is not practical to visualize and display all molecular configuration samples. For this reason, the information processing device 100 selects, from among a large number of molecular configuration samples, a representative molecular configuration sample that is useful for understanding the statistical information. However, among the various molecular configuration samples generated by molecular simulation, some do not adequately represent the trends in intermolecular interactions indicated by the statistical information. For example, the molecular configuration indicated by a certain molecular configuration sample may have occurred by chance and may not reflect frequently occurring intermolecular interactions. Visualizing a molecular configuration sample that does not reflect frequently occurring intermolecular interactions may actually make it more difficult to understand the statistical information. Therefore, it is important for the information processing device 100 to determine how to select a representative molecular configuration sample.
[0055] As an example of a method for selecting a molecular configuration sample, it is possible to select the molecular configuration sample that was generated last. However, the appropriateness of the last molecular configuration sample depends on the simulation time setting in the molecular dynamics simulation and the number of iterations setting in the Monte Carlo simulation. Therefore, the last molecular configuration sample does not necessarily reflect the intermolecular interactions that occur frequently.
[0056] Another possible method for selecting molecular configuration samples is to calculate the root mean square deviation (RMSD) between molecular configuration samples and select the molecular configuration sample located at the center in the RMSD space. RMSD is calculated by calculating the difference in the position coordinates of each atom between two molecular configuration samples and then finding the root mean square of the differences for multiple atoms. The molecular configuration sample selected by this method represents the average molecular configuration among the molecular configurations that appeared in the molecular simulation. However, the central molecular configuration sample is easily influenced by a small number of exceptional molecular configuration samples and may deviate from the ideal molecular configuration sample that reflects frequently occurring intermolecular interactions.
[0057] Another possible method for selecting molecular configuration samples is to calculate a score for each molecular configuration sample using a pre-prepared score function and select molecular configuration samples with high scores. The score could be, for example, a docking score, which indicates the general binding affinity between molecules. However, the pre-prepared score function objectively and uniquely calculates a score based on the molecular structure and intermolecular distance, and does not take into account the simulation conditions of this molecular simulation, such as the presence of water molecules. Therefore, it may not be possible to calculate appropriate scores for multiple molecular configuration samples that are consistent with the statistical information.
[0058] Therefore, the information processing device 100 dynamically generates a score function from multiple molecular arrangement samples generated by molecular simulation, calculates the score of each molecular arrangement sample using the generated score function, and selects molecular arrangement samples with high scores.
[0059] FIG. 5 is a diagram showing examples of interatomic distances and cutoff distances. First, the information processing device 100 generates a score function that calculates the weight of pairs of atoms belonging to different molecules. Pairs are extracted comprehensively. Pairs include pairs of atoms of a protein and a water molecule, pairs of atoms of a low molecular weight compound and a water molecule, pairs of atoms of a protein and an atom of a low molecular weight compound, etc. However, hydrogen atoms are excluded from the atoms that form pairs.
[0060] The information processing device 100 generates a score function shown in Equation (3) using all molecular configuration samples generated by molecular simulation. In Equation (3), M is the number of molecular configuration samples, R n ij is the distance between atoms i and j in molecular configuration sample n, R0 is the cutoff distance, F R0 ij is the weight of atoms i and j calculated based on the cutoff distance R0. The cutoff distance R0 is a threshold indicating the range in which intermolecular interactions are considered. The cutoff distance R0 corresponds to parameter 16 in the first embodiment.
[0061]
number
[0062] f is a potential function. The potential function f outputs a numerical value indicating the strength of the interaction between atoms. The output of the potential function f depends on the interatomic distance and the elements of the two atoms. The smaller the interatomic distance, the larger the value output by the potential function f. Also, if the two atoms are carbon atoms, the potential function f outputs a relatively large value. However, the potential function f may also be a function that outputs a constant (for example, 1). δ is a delta function. The delta function δ outputs 0 if its argument is negative, and outputs 1 if its argument is non-negative (greater than or equal to 0).
[0063] Therefore, the weight calculated by the score function for the pair of atoms i and j is determined from the cutoff distance R0 and the distance between atoms i and j in M molecular configuration samples. The weight for the pair of atoms i and j is the average potential obtained by summing up the potentials for the molecular configuration samples in which the distance between atoms i and j is equal to or less than the cutoff distance R0 among the M molecular configuration samples, and dividing this sum by the number of molecular configuration samples M. In particular, if the potential function f always outputs 1, the weight for the pair of atoms i and j indicates the proportion of the molecular configuration samples in which the distance between atoms i and j is equal to or less than the cutoff distance R0 among the M molecular configuration samples.
[0064] In the molecular configuration sample 41, atom 42 corresponds to atom i, and atom 43 corresponds to atom j. The interatomic distance is the distance between the centers of the atoms. The distance R between atoms 42 and 43 n ij is equal to or less than the cutoff distance R0. Therefore, the value of the potential function f in the molecular configuration sample 41 is reflected in the weight of the pair of atoms 42 and 43.
[0065] However, the cutoff distance R0 is unknown when the score function is generated from the molecular configuration sample, and is optimized by a method described below. Therefore, the information processing device 100 defines the score function so as to output a weight according to the cutoff distance R0. For example, the information processing device 100 calculates weights corresponding to various cutoff distances R0, such as a weight when R0=3.2, a weight when R0=3.3, a weight when R0=3.4, and so on.
[0066] Next, the information processing device 100 optimizes the cutoff distance R0. At this time, the information processing device 100 calculates an evaluation value s shown in Equation (4). The evaluation value s for the cutoff distance R0 is calculated by <ij>The evaluation value s is calculated by summing up the weights for each and dividing by the cube of the cutoff distance R0. However, the denominator may be the volume of a sphere with a radius equal to the cutoff distance R0. When the potential function f always outputs 1, the evaluation value s is proportional to the average atomic density within the sphere of radius R0. The information processing device 100 calculates the derivative Δs of the evaluation value s with respect to the cutoff distance R0, as shown in Equation (5).
[0067]
number
[0068]
number
[0069] FIG. 6 is a graph showing an example of the derivative of the average atomic density. Graph 44 shows the relationship between cutoff distance R0 and derivative Δs of the evaluation value. The horizontal axis of graph 44 represents cutoff distance R0 (unit: angstroms). The vertical axis of graph 44 represents derivative Δs of the evaluation value corresponding to cutoff distance R0.
[0070] When the cutoff distance R0 is gradually increased, the differential Δs increases when the cutoff distance R0 is small because there are many atoms with stable interactions, i.e., atoms with a high probability of being nearby. On the other hand, when the cutoff distance R0 is sufficiently large, the differential Δs decreases because there are many atoms with no stable interactions, i.e., atoms with a low probability of being nearby. Therefore, the information processing device 100 selects the cutoff distance R0 that maximizes the differential Δs of the evaluation value as the optimal value so that the score function can determine whether the pair of atoms has a stable interaction. In the example of graph 44, the information processing device 100 changes the cutoff distance R0 from 3.2 to 5.0 in 0.1 increments and determines R0 = 3.9 as the optimal cutoff distance R0.
[0071] Next, for each of the multiple molecular configuration samples, a score is calculated as shown in Equation (6). In Equation (6), R0 is the cutoff distance optimized by the above method, and S n is the score for molecular configuration sample n. n is the set of all pairs of atoms that satisfy the above conditions. <ij>The information processing device 100 calculates the score S n Sort multiple molecular configuration samples in descending order of , and score S n A predetermined number of one or more molecular configuration samples are selected from those with the highest values, and the selected molecular configuration samples are visualized.
[0072]
number
[0073] In the molecular configuration sample selection method of the second embodiment, a high weight is assigned to pairs of atoms that are frequently close to each other across all molecular configuration samples. Furthermore, a molecular configuration sample in which pairs of atoms with high weights are actually close to each other receives a high score. Therefore, a molecular configuration sample that reflects frequently occurring intermolecular interactions is selected. Furthermore, a molecular configuration sample to be visualized is uniquely determined based on the multiple generated molecular configuration samples.
[0074] The cutoff distance is automatically optimized and does not depend on user input. Using an inappropriate cutoff distance may result in an inappropriate score function being generated, potentially resulting in the selection of an inappropriate molecular configuration sample. If the cutoff distance is too small, the range of partner atoms considered for interaction may be too narrow, resulting in a score function that does not properly evaluate whether an atom pair has a stable interaction. On the other hand, if the cutoff distance is too large, the score function may take into account interactions with distant partner atoms that do not contribute to the stabilization of the molecular configuration, resulting in noise in the score function weights.
[0075] Furthermore, the appropriate cutoff distance varies depending on the molecular simulation. Different types of proteins and low molecular weight compounds result in different types of interacting atom pairs and different relative positions of those atom pairs. Furthermore, different initial positions and initial orientations of low molecular weight compounds can result in the low molecular weight compound binding to different parts of the protein and binding to the protein in different orientations. In contrast, the molecular configuration sample selection method of the second embodiment calculates a cutoff distance that is appropriate for the molecular simulation at that time.
[0076] FIG. 7 is a diagram showing a comparison between the experimental results and the simulation results. Molecular configuration sample 51 represents the correct molecular configuration determined by chemical experiments and X-ray crystal structure analysis. Molecular configuration sample 52 is the molecular configuration sample with the lowest score among the molecular configuration samples generated by molecular simulation. Molecular configuration sample 53 is the molecular configuration sample with the highest score.
[0077] Amino acid residue list 54 shows amino acid residues that are within a cutoff distance R0 from the low molecular weight compound in molecular configuration sample 51. Amino acid residue list 55 shows amino acid residues that are within a cutoff distance R0 from the low molecular weight compound in molecular configuration sample 52. Amino acid residue list 56 shows amino acid residues that are within a cutoff distance R0 from the low molecular weight compound in molecular configuration sample 53.
[0078] An amino acid residue is a part of a protein that corresponds to one amino acid unit. Amino acid residue lists 54, 55, and 56 include the position number indicating the position on the protein and the name of the amino acid residue. TRP is tryptophan, PRO is proline, PHE is phenylalanine, GLN is glutamine, VAL is valine, LEU is leucine, TYR is tyrosine, CYS is cysteine, ASN is asparagine, and ILE is isoleucine.
[0079] Amino acid residue list 54 includes 11 amino acid residues. Amino acid residue list 55 includes 3 amino acid residues out of the 11 amino acid residues shown in amino acid residue list 54. Amino acid residue list 56 includes 10 amino acid residues out of the 11 amino acid residues shown in amino acid residue list 54. Therefore, molecular configuration sample 53 with the highest score reproduces the experimental measurement results with high accuracy.
[0080] Next, the functions and processing procedures of the information processing device 100 will be described. FIG. 8 is a block diagram illustrating an example of functions of the information processing device. The information processing device 100 has a molecular configuration storage unit 121, a statistical information storage unit 122, a molecular simulation unit 123, a score calculation unit 124, and a visualization unit 125. The molecular configuration storage unit 121 and the statistical information storage unit 122 are implemented using, for example, a RAM 102 or an HDD 103. The molecular simulation unit 123, the score calculation unit 124, and the visualization unit 125 are implemented using, for example, a CPU 101 and a program.
[0081] The molecular configuration storage unit 121 stores definitions of molecules to be placed in the simulation space. The molecular configuration storage unit 121 also stores a plurality of molecular configuration samples generated by molecular simulation. Each molecular configuration sample includes the position coordinates of each of a plurality of atoms. In the case of molecular dynamics simulation, the molecular configuration storage unit 121 stores a molecular configuration sample of the initial configuration and a molecular configuration sample at each time. In the case of Monte Carlo simulation, the molecular configuration storage unit 121 stores randomly generated molecular configuration samples.
[0082] The statistical information storage unit 122 stores statistical information of physical property values generated by molecular simulation. For example, the statistical information storage unit 122 stores at least one of the interaction energy and the average energy at each time shown in FIG. 4 as statistical information.
[0083] The molecular simulation unit 123 executes a molecular simulation based on the molecular definition stored in the molecular configuration storage unit 121. For example, the molecular simulation unit 123 executes a molecular dynamics simulation or a Monte Carlo simulation. The molecular simulation unit 123 stores a plurality of molecular configuration samples in the molecular configuration storage unit 121. The molecular simulation unit 123 also generates statistical information of physical property values from the plurality of molecular configuration samples and stores the statistical information in the statistical information storage unit 122.
[0084] The score calculation unit 124 generates a score function using a plurality of molecular configuration samples stored in the molecular configuration storage unit 121. The score calculation unit 124 optimizes the cutoff distance so as to maximize the derivative of the evaluation value, and calculates the score of each molecular configuration sample using the score function and the optimized cutoff distance.
[0085] The visualization unit 125 selects one or a predetermined number of molecular arrangement samples having higher scores from among the plurality of molecular arrangement samples. The visualization unit 125 generates image data that visualizes the relative arrangement of molecules indicated by the selected molecular arrangement sample. The visualization unit 125 outputs the generated image data in association with the statistical information stored in the statistical information storage unit 122. The visualization unit 125 may display the statistical information and image data on the display device 111, may store them in nonvolatile storage, or may transmit them to another information processing device.
[0086] FIG. 9 shows an example of a molecule list and a molecule configuration table. The molecule list 126 is stored in the molecular configuration storage unit 121. The molecule list 126 associates molecule IDs, atom IDs, and elements. A molecule ID is an identifier that identifies a molecule such as a protein, a low molecular weight compound, or a water molecule. An atom ID is an identifier that identifies an atom that forms a molecule. An element is a type of atom such as hydrogen H, oxygen O, or carbon C. The molecular configuration storage unit 121 also stores information indicating the connection relationships between multiple atoms within the same molecule.
[0087] The molecular configuration table 127 is stored in the molecular configuration storage unit 121. The molecular configuration table 127 associates a sample ID, an atom ID, and a position coordinate. The sample ID is an identifier for identifying a molecular configuration sample. The atom ID is an identifier for identifying an atom. The position coordinate is a three-dimensional coordinate indicating the position of an atom in the molecular configuration sample.
[0088] FIG. 10 is a diagram illustrating an example of a score function table. The score function table 128 is generated by the score calculation unit 124 and stored in the RAM 102. The score function table 128 indicates the score function before the cutoff distance is optimized. The score function table 128 stores weights calculated based on each of a plurality of cutoff distances, in association with two atom IDs indicating an atom pair.
[0089] For example, the score function table 128 lists multiple weights for a pair of the second atom and the fourth atom, such as a weight when R0=3.2, a weight when R0=3.3, a weight when R0=3.4, etc. Similarly, the score function table 128 lists multiple weights for a pair of the second atom and the sixth atom, such as a weight when R0=3.2, a weight when R0=3.3, a weight when R0=3.4, etc.
[0090] FIG. 11 is a flowchart showing an example of a procedure for visualizing a molecular arrangement. (S10) The molecular simulation unit 123 generates M molecular configuration samples in which at least some of the atomic positions are different through molecular simulation. For example, the molecular simulation unit 123 generates M molecular configuration samples corresponding to different times through molecular dynamics simulation. Also, for example, the molecular simulation unit 123 randomly generates M molecular configuration samples through Monte Carlo simulation.
[0091] (S11) The molecular simulation unit 123 generates statistical information of physical property values from the M molecular configuration samples generated in step S10. For example, the molecular simulation unit 123 calculates the interaction energy between a protein and a low molecular weight compound.
[0092] (S12) The score calculation unit 124 selects a pair of atoms (atoms i, j) other than hydrogen atoms that belong to different molecules. The score calculation unit 124 calculates the distance between atoms i and j for each of the M molecular configuration samples.
[0093] (S13) The score calculation unit 124 substitutes the M distances calculated in step S12 into Equation (3) to express the weight of the pair of atoms i and j using the cutoff distance. For example, the score calculation unit 124 calculates the weight corresponding to each candidate cutoff distance.
[0094] (S14) The score calculation unit 124 determines whether all atom pairs have been selected in step S12. If all atom pairs have been selected, the process proceeds to step S15, and if there are any unselected atom pairs, the process returns to step S12.
[0095] (S15) The score calculation unit 124 calculates an evaluation value indicating the average atomic density by dividing the sum of the weights of all atom pairs by the cube of the cutoff distance. The score calculation unit 124 calculates the derivative of the evaluation value while increasing the cutoff distance by a fixed amount. The derivative of the evaluation value is simply calculated, for example, by dividing the difference between the evaluation value corresponding to a certain cutoff distance and the evaluation value corresponding to the previous cutoff distance by the increase in the cutoff distance.
[0096] (S16) The score calculation unit 124 selects, as the optimum value, the cutoff distance at which the derivative of the evaluation value calculated in step S15 is maximized. (S17) The score calculation unit 124 applies the cutoff distance selected in step S16 to the score function. The score calculation unit 124 calculates the score of each of the M molecular arrangement samples using the score function according to Equation (6).
[0097] (S18) The visualization unit 125 extracts the molecular configuration sample with the highest score. However, the visualization unit 125 may extract two or more molecular configuration samples with the highest scores. (S19) The visualization unit 125 generates image data that visualizes the molecular arrangement sample extracted in step S 18. For example, the visualization unit 125 generates image data that represents atoms as points, circles, or spheres according to the position coordinates indicated by the extracted molecular arrangement sample.
[0098] (S20) The visualization unit 125 displays the statistical information generated in step S11 and the image data generated in step S19 in association with each other. As described above, the information processing device 100 of the second embodiment selects and visualizes some molecular configuration samples from a large number of molecular configuration samples generated by molecular simulation, and displays them together with statistical information. This makes it easier for the user to understand the statistical information regarding intermolecular interactions. Furthermore, the information processing device 100 calculates a score for each of the multiple molecular configuration samples and selects molecular configuration samples with high scores. This allows the selection of molecular configuration samples that are consistent with the statistical information, compared to when the last generated molecular configuration sample or a randomly selected molecular configuration sample is visualized.
[0099] Furthermore, the information processing device 100 generates a score function and optimizes the cutoff distance using multiple molecular configuration samples generated in the current molecular simulation. As a result, compared to using a fixed score function or a fixed cutoff distance, an appropriate score function suited to the current molecular simulation is generated, and atom pairs contributing to intermolecular interactions are appropriately identified. Furthermore, the information processing device 100 calculates the score of the molecular configuration sample by summing the weights of atom pairs whose distance is equal to or less than the cutoff distance. As a result, a molecular configuration sample reflecting frequently occurring intermolecular interactions is selected.
[0100] Furthermore, the information processing device 100 calculates an evaluation value indicating the average atomic density by dividing the sum of the weights of various atom pairs by the unit space size corresponding to the cutoff distance, and determines the cutoff distance so that the derivative of the evaluation value is maximized. As a result, the cutoff distance is determined so that the score function covers a range in which there are many atoms having stable interactions, and atom pairs that contribute to frequently occurring intermolecular interactions are appropriately evaluated.
[0101] The foregoing merely illustrates the principles of the present invention. Further, since numerous modifications and changes will be apparent to those skilled in the art, the present invention is not limited to the exact construction and application shown and described above, and all corresponding modifications and equivalents are deemed to be within the scope of the present invention as defined by the appended claims and their equivalents. [Explanation of symbols]
[0102] 10. Information processing equipment 11 Storage section 12 Processing section 13,14 Molecular configuration sample 15 Score Functions 16 Parameters 17 Priority< / ij> < / ij>
Claims
1. Acquire a plurality of molecular configuration samples each including position information of a plurality of atoms; generating a score function that calculates a weight for a pair of atoms belonging to different molecules among the plurality of atoms, the score function including a parameter indicating a distance threshold, based on the distances between the pair of atoms in the plurality of molecular configuration samples; determining a specific value to be set for the parameter based on a change in weight calculated by the score function in response to a change in the value of the parameter; determining priorities of the plurality of molecular configuration samples based on specific weights calculated by the score function using the specific values and distances between the pairs of atoms in each of the plurality of molecular configuration samples; A structure extraction program that executes processing on a computer.
2. the score function calculates a weight according to the proportion of molecular configuration samples in which the distance between the pair of atoms is equal to or less than the parameter value, among the plurality of molecular configuration samples; The structure extraction program according to claim 1.
3. In determining the specific value, a change in an evaluation value obtained by dividing the weight calculated by the score function by a space size corresponding to the parameter value is monitored in response to a change in the parameter value, and the specific value is determined so that a derivative of the evaluation value is maximized. The structure extraction program according to claim 1.
4. In determining the priority order, a score is calculated for each of the plurality of molecular configuration samples based on whether or not the distance between the pair of atoms is equal to or less than the specific value and the specific weight, and the priority order is determined based on the score. The structure extraction program according to claim 1.
5. generating image data that visualizes position information of the plurality of atoms included in the molecular configuration sample with the highest priority among the plurality of molecular configuration samples, and outputting the image data; The structure extraction program according to claim 1.
6. Acquire a plurality of molecular configuration samples each including position information of a plurality of atoms; generating a score function that calculates a weight for a pair of atoms belonging to different molecules among the plurality of atoms, the score function including a parameter indicating a distance threshold, based on the distances between the pair of atoms in the plurality of molecular configuration samples; determining a specific value to be set for the parameter based on a change in weight calculated by the score function in response to a change in the value of the parameter; determining priorities of the plurality of molecular configuration samples based on specific weights calculated by the score function using the specific values and distances between the pairs of atoms in each of the plurality of molecular configuration samples; A structure extraction method in which processing is performed by a computer.
7. a storage unit that stores a plurality of molecular configuration samples each including position information of a plurality of atoms; a processing unit that generates a score function, which is a function for calculating weights for pairs of atoms among the plurality of atoms that belong to different molecules, based on the distances between the pairs of atoms in the plurality of molecular configuration samples, the score function including a parameter indicating a distance threshold, determines a specific value to be set for the parameter based on a change in the weight calculated by the score function in response to a change in the value of the parameter, and determines the priority of the plurality of molecular configuration samples based on the specific weight calculated by the score function using the specific value and the distances between the pairs of atoms in each of the plurality of molecular configuration samples; An information processing device having the above.
Citation Information
Patent Citations
Method for displaying molecule
JP1989084392A
Device and method of simulation and recording medium storing simulation program
JP2006190234A
Device, method, and program for computing interaction between protein and chemical compound
JP2007219789A
Simulation system
JP2008158781A
Particle behavior analysis device and program
JP2008269443A